Back to Explore
NewsVibeFix 编辑部Updated Oct 9, 2026

One Prompt, Six Hours: Opus 5.5 and GPT-6 Astra Recreate Calvino's Invisible Cities in a Vibe Coding Showdown

One prompt, six unsupervised hours: GPT-6 Astra finished a three.js visualization of Calvino's 55 Invisible Cities in 53 minutes for $10, while Claude Opus 5.5 took 85 minutes, 6 parallel subagents and $74 — and declared a 'jump' for design tasks. A field report on vibe coding's limit case, model selection, and the token ledger.

Artistic 3D visualization of imaginary floating cities at night, glowing nodes and light trails forming an urban constellation

This is the most extreme public vibe coding experiment to date: a human says exactly one sentence, and for the next six hours the machine works alone — no check-ins, no clarifying questions, no changing requirements. In early October, developer Piotr Migdał published a dual-model showdown on the Quesma company blog — the same prompt handed to GPT-6 Astra running on Codex and to Claude Opus 5.5 running on Claude Code, with one task: turn the 55 imaginary cities of Italo Calvino's Invisible Cities into a three.js visualization. The experiment picked up 274 points and 144 comments on Hacker News, and dev.to featured it in its October 8, 18:00 digest. One note on sourcing: the Quesma blog page itself carries no publication date stamp; the October 8/9 window is inferred from the dev.to digest and the HN submission record, and this article does not present it as a date printed on the original page.

One Prompt, Six Hours

The entire input to the experiment was this single prompt, verbatim:

Make a three.js (pnpm) visualization of all Invisible Cities by Italo Calvino. Don't ask questions, it is a one-shot task. You have 6h of work, use it until it becomes a masterpiece.

The operative phrase is "Don't ask questions." In a normal agent workflow, the model stops to ask when it hits ambiguity: where does the data come from? What visual style? How deep should the interactivity go? Migdał shut that door — no questions allowed, one shot, six hours to spend freely, with "masterpiece" as the target.

The choice of Invisible Cities as the test is deliberate. This is not a "turn this CSV into a bar chart" task with a right answer. Calvino's 55 cities are pure literary imagination — no coordinates, no satellite imagery, not even a single canonical answer to what they look like. The model first has to "read" the textual imagery, then translate it into 3D space, color, motion, and interaction. At least half the job is aesthetic judgment, half is engineering — sitting exactly on the fault line between "design" and "puzzle-solving," which makes it an ideal instrument for comparing the temperaments of two models.

Two Report Cards: 53 Minutes vs. 85 Minutes

First, GPT-6 Astra. It delivered in 53 minutes at a token cost of about $10. Fast, cheap, one clean run.

Then Claude Opus 5.5. It took 1 hour 25 minutes, dispatched 6 parallel subagents totaling roughly 7 agent-hours, at a token cost of about $74. The most telling detail: it never used up its budget — in the author's words, the model's own verdict was "I used roughly half of the six hours." It finished early, on its own call.

  • GPT-6 Astra (Codex): 53 minutes, ~$10 in tokens
  • Claude Opus 5.5 (Claude Code): 1 hour 25 minutes, 6 parallel subagents totaling ~7 agent-hours, ~$74 in tokens, finished ahead of schedule

Migdał's conclusion is restrained, but every word carries weight: Astra is "a jump" on puzzle-like tasks; Opus 5.5 is "a jump" on design-like tasks. Note what he did not say — he didn't declare a winner. He described a division of labor. This isn't an arena; it's a org chart.

From Pairing to Async: Vibe Coding's Limit Case

What deserves discussion here isn't who won, but how the work happened. For two years, "vibe coding" has really meant human-machine pairing: the human states intent, the model writes code, the human reviews, the human refines the prompt, repeat. Human attention stays in the loop at all times, acting as quality control on the line — step away and things fall over.

Migdał's experiment removed the human from the loop. Across six hours, humanity's total output was that one sentence at the top. The model planned, wrote, and revised on its own — Opus 5.5 even decided for itself that half the allotted time was enough and clocked out early. This is no longer pairing; it's async work: you drop a task into a black box, go do something else, and come back to review.

The precondition for this mode is that a model doesn't drift or collapse during long, unsupervised runs. Six hours and 7 agent-hours of autonomous work would have been unthinkable a year ago: agents of that era would spiral into loops, break their own code, or stall waiting for a human decision after thirty minutes. The "finished early" detail says more than any benchmark score — it wasn't grinding out the full six hours, it was converging on its own judgment that the work was done. Self-convergence in long-horizon tasks is the real ticket to human-machine async. With that ticket, vibe coding stops being "efficient pairing" and becomes "outsourced labor": your time is no longer the bottleneck — the model's autonomy is.

The Design-Task Watershed: Match the Model to the Task

Translated into a selection guide, the "a jump" verdict reads: puzzles go to Astra, design goes to Opus 5.5.

Puzzle-like tasks — long chains of logic, relatively unique answers, demanding rigorous reasoning: writing a parser, fixing a concurrency bug, solving an algorithms problem. Astra is plausibly the faster, cheaper answer. Design-like tasks — heavy on aesthetic judgment, no standard answer, full of trade-offs in open space: turning 55 literary cities into an interactive 3D visualization. There, the Opus 5.5 premium may well be worth it.

The phrase "a jump" also hints at a larger backdrop: design-like tasks used to be every large model's weak spot. They could all do logic; asked "does it look good, is it fun," they'd flinch. That gap is now closing. Once design stops being a model weakness, "a decent visual piece from a single sentence" graduates from demo trick to routine operation — and vibe coding's territory expands by a wide margin.

Notably, this isn't Opus 5.5's first appearance in a "one sentence, long task" setting. Earlier, Ryan Sael used Opus 5.5 to build an interactive lens lab in one shot: 1 hour 26 minutes, $25.66 in tokens. The timescales of the two experiments are strikingly consistent — just over an hour each, costs in the tens of dollars. "One sentence + one hour + tens of dollars" is crystallizing into the canonical profile of this task class: not a one-off fluke, but a reproducible working pattern.

$10 vs. $74: A Ledger You Have to Keep

Fifty-three minutes / $10 against 85 minutes / $74: more than 7x the cost, 1.6x the time. On raw numbers, Astra wins outright. But that's not how you read this ledger.

First, what Opus burned was parallel compute: 6 subagents pushing forward simultaneously, 7 agent-hours in total, yet only half an hour more on the wall clock. You're not buying "slower" — you're buying "wider": six parallel branches of exploration converging into one artifact. That's rational for design work: good design is divergence first, convergence second. A single thread grinding away is more likely to get stuck in a local optimum.

Second, cost has to be weighed against output. $74 bought a complete unsupervised delivery: from a prompt to a live interactive demo (playable at p.migdal.pl/invisible-cities-opus-5.5) to fully open-sourced code (github.com/stared/invisible-cities-opus-5.5). If this were a small "design + front-end" freelance gig, $74 wouldn't buy you one hour of human labor. The vibe coding cost equation always has tokens in the numerator and "how many human-hours did this replace" in the denominator — computed that way, $74 is absurdly cheap.

And yet $74 is not pocket change either. At that unit price, ten such tasks a day is $740, over $20,000 a month. That may be this blog post's most practical contribution: it puts $10 and $74 side by side on the same page and forces you to take budgeting seriously. When do you send in the $10 quick-draw, and when is it worth launching the $74 parallel fleet? You need a doctrine — and task-type triage (puzzle vs. design) is rule number one. As agent tasks become a daily expense, the "token budget" will join the "cloud bill" as required reading for every team.

A Cooler Take: A Demo Is Not Production Code

For all that, this kind of experiment has a built-in limitation: its acceptance criterion is "masterpiece," not "maintainable," "testable," or "shippable." A visualization demo has enormous fault tolerance — nobody notices a missing city, nobody complains about a dropped frame, ugly code is fine as long as it runs. Production code is judged by a completely different rubric: edge cases, error handling, long-term maintainability — every one of them things a demo can hand-wave past and a 3 a.m. pager cannot.

So don't extrapolate "six unsupervised hours" into "a production system in six hours." What the experiment proves is a ceiling — the autonomous capability of models on open-ended creative work. What it cannot prove is a floor — whether, under tight engineering constraints, a model will sacrifice correctness to "ship." Between demo and production still stand three gates — code review, testing, and the production incident — any of which can send "unsupervised" back to "supervised."

Also, the sample size is 1: one prompt, two models, one run each. Was there luck in Astra's 53 minutes? Does Opus's 6-subagent strategy work every time? Would the conclusion hold on a different problem? Unknown. The value of a controlled experiment is in raising good questions, not in settling them. Use it as a selection reference, not a procurement justification — for that, it's still too early.

Closing: The Sentence Is Becoming the New Compiler

Zoom out, and the experiment marks a turning point: the prompt is graduating from "conversation opener" to "task compiler." You used to write prompts to spar with a model across many rounds; now you can write one prompt and go do something else — the model plans, executes, and wraps up on its own, and you come back to review. From human-machine pairing to human-machine async, what's saved isn't just typing time but attention itself.

And the double "a jump" tells us the capability map is differentiating: there is no universal champion, only fit between model and task type. The 2026 vibe coder should hold two cards — the quick-draw for puzzles, the fleet for design — plus a ledger that knows which to play when. What Migdał showed us with 55 imaginary cities is roughly this: on the right task, give the right model one sentence and six hours, and it may hand you back something at masterpiece level — or at minimum, a reason to redo your math.

Primary source: the original Quesma blog post (no date stamp on the page; the time window is inferred from the dev.to October 8 digest and the HN October 9 submission record). The interactive demo and full source are linked above.

Sources

Browse projectsPublish your project

Related articles

Docker Agent news cover: developer terminal screen running Docker containers, real photo
News
Docker Open-Sources Docker Agent: Define an Agent in YAML, Run It Like a Container

On Oct 8, Docker's docker CLI plugin Docker Agent hit the HN front page (290 points / 133 comments). It defines agents in declarative YAML (agent.yaml) — 'no code required' — with container-style commands. Model-agnostic (7 providers), native MCP tools, built-in think/todo/memory, full RAG stack, multi-agent delegation. Killer move: agents push/pull to any OCI registry like images, making definitions versionable, PR-reviewable artifacts. Repo dates to Sep 2025: 10,735 commits, 263 releases.

Open-source ProjectsAI CodingDeveloper Workflow
Terminal output showing tsc-rs type-checking benchmark results
News
$400K in Tokens Burned, Zero Lines Read: AI Ported the TypeScript Compiler to Rust in Two Weeks

Theo Browne had LLMs port the TypeScript 7 compiler to Rust: 181,711 tests green, 13x faster than tsc 6 on VS Code. The real story is the bill — $400K of Codex tokens got stuck at 84%, then Claude Opus 5.5 finished it in two weeks. Plus the trust question nobody can dodge: the author has never read a line of the code.

Open-source ProjectsTrending ProjectsAI Coding
A laptop screen showing a website signup page inside a browser
News
ChatGPT Sites Hits HN's Front Page: Prompt-to-Website — Toy or Productivity?

On October 3, 'Sites in ChatGPT' hit the HN front page with ~209 points and 218 comments. Not a launch — a reckoning: is prompt-to-URL a toy, a prototype host, or a productivity tool? The four debates, the doc-backed facts (D1/R2, sign-in, custom domains), and three verdicts for vibe coders.

AI CodingProduct LaunchIndie Development