Back to Explore
NewsVibeFix 编辑部Updated Oct 10, 2026

Simon Willison Built a Blog Feature "Almost Entirely Using My Voice" — While Cooking Dinner

Simon Willison shipped a Newsletters index page for his blog almost entirely by voice — chatting with the Codex tab in the ChatGPT desktop app while cooking dinner in his kitchen, barely touching the keyboard. When a top-tier practitioner starts coding with his mouth, voice + agents stop being a gimmick and become real productivity. A teardown of his playbook, where this workflow breaks, and the minimal setup to copy him.

Concept illustration: a laptop on a kitchen counter in voice conversation with a coding agent, code on screen, cooking pots beside it

On October 9, 2026, Simon Willison published a 1,294-word post on his blog titled "A new feature for my blog, built using my voice." The title tells the whole story: he shipped a new feature for his blog — a Newsletters index page collecting every newsletter he has ever sent, both the free weekly Substack and the paid monthly sponsors-only updates. And he built the entire feature, in his words, "almost entirely using my voice" — while cooking dinner in his kitchen.

This deserves a full article not because the feature was hard — quite the opposite, it was deliberately simple. It matters because of who and how: Simon Willison, creator of Datasette and co-creator of Django, one of the most meticulous documenters of LLM practice working today. Every workflow post on his blog reads like a preview of where the developer community will be six months from now. When someone like that starts "writing code with his mouth" — and bothers to write the whole process down — this stops being a voice-assistant marketing gimmick and becomes a productivity shift worth taking apart.

What he actually built: an index page, in the time it takes to cook dinner

First, the feature. The new Newsletters page (simonwillison.net/newsletters/) does a tidying-up job: indexing newsletters that lived in two different places — the free weekly Substack (simonw.substack.com, 70,000+ subscribers) and the monthly updates for GitHub Sponsors. For readers, it is finally one place to find back issues; for the blog itself, it is a standard piece of content-asset housekeeping.

Then the scope. Willison breaks the work down plainly in the post: a standard small Django feature — a new model, a migration, some view code, templates, plus two import functions to populate the database from external sources. Sounds simple, and that is exactly the point: he picked a task the agent was guaranteed to nail, to validate the workflow — not a hard problem to wrestle by voice. Trying a new posture on an easy win first: that is veteran judgment.

Reconstructing his playbook: he typed exactly one line

The most valuable part of the post is how he documented the procedure like a lab notebook, down to which button to click:

  • The tool: the ChatGPT desktop app, Codex tab, voice conversation mode, running against a local development environment — his blog's codebase, simonwillisonblog, checked out on the laptop, with the agent working directly on it.
  • The opening move: he typed exactly one serious line — asking the agent to start the dev server and open the preview in a browser. The brilliance of this step is that it established a visible acceptance loop: from then on he could ask the agent to "show me the new pages" and track progress with his eyes.
  • The button that matters: then he clicked the "Start new voice chat" button — stressing that this is not the microphone button, but the one to its right. Only a detail obsessive like Willison would write that distinction down, and it is precisely the kind of trap-clearing future copiers need.
  • The session: laptop set up in the kitchen, talking to it while cooking. The model under the hood: GPT-6 Astra High.
  • The receipts: he even pasted an extract of the voice transcript captured by Codex — publishing the raw dictation log of "mouth-driven development" for anyone to audit, theater or real work.
Concept illustration: a laptop on a kitchen counter showing a voice conversation with a coding agent, cooking pots beside itWillison's setup, schematically: laptop in the kitchen, voice chat driving a local coding agent, browser preview as the acceptance loop

Note the division of labor in this chain: the mouth issues instructions, the agent writes code and runs services, the eyes check the preview. No step is "wishing into the void" — everything has a landing point. That is why this post is credible: not another mystical "I had AI build a blog feature" story, but a controlled experiment with acceptance criteria and a transcript.

Why now: four preconditions for voice vibe coding

Voice programming is not a new idea. A decade ago people were shouting "open file" at IDEs, and it was generally a disaster. Willison pulled it off not because he has a great voice, but because four preconditions all landed in 2026:

First, the agent loop grew up. Coding agents are no longer toys that generate one snippet at a time; they are closed loops that start servers, open browsers, and look at rendered pages themselves. Willison's first step — having the agent boot the dev server and open a preview — effectively outsourced the "hands" to the agent and kept "acceptance" for himself. Voice is good at saying what to do and bad at line-by-line checking of how; the agent loop covers exactly the second half. Without it, voice would just be a fancy dictation machine.

Second, context got long enough. Voice is high-entropy input: spoken language rambles, repeats, self-corrects — "uh, I mean, that, that index page…" Only a long-context model that remembers the whole conversation and can re-read the repo at will can turn "you know what I mean" into a correct migration. In the short-context era, the information loss of voice coding was fatal.

Third, transcription quality cleared the bar. One mistranscribed word can mean one wrong variable name. That Willison trusted transcription end-to-end tells you this link in the chain is now good enough to stop worrying about — two years ago it still needed proofreading.

Fourth, and most important: the agent absorbs the precision loss. Keyboard input is high-precision, low-bandwidth; voice is high-bandwidth, low-precision. This road was blocked before because code does not tolerate precision loss — one wrong character and nothing compiles. Now the agent sits in the middle as a translation layer, compiling fuzzy spoken intent into exact code. The weakness of the input modality got patched by the intelligence of the execution layer.

So "coding with your mouth" is not a single-point victory for speech technology; it is the input end coming unbound once the whole agent stack matured. That Willison chose this moment to do it is itself a judgment call.

"I built the feature almost entirely using my voice, chatting away to my laptop while I cooked dinner."

— Simon Willison, "A new feature for my blog, built using my voice," October 9, 2026

The boundaries: what fits "cooking-dinner development" and what doesn't

Don't throw away your keyboard just yet. Willison's experiment is actually a perfect template of when this works — invert it, and you get when not to try:

Tasks that fit look like this:

  • The requirements are already formed in your head. He says he had "a pretty good idea of what I wanted to build." Voice is for expressing known intent, not for designing on the fly — you can't really say to a pot, "let me think about this architecture…"
  • Small scope, cheap mistakes. One page, one migration; a botched attempt rolls back cheaply. Voice's high bandwidth is an advantage here: describing a CRUD page out loud is far faster than typing it.
  • A short, visual acceptance loop. A glance at the browser is enough; no 500-line logs to read. Note the order in Willison's playbook: preview loop first, voice second. Never reversed.
  • You know the codebase cold. This is his own blog, a Django project — he knows what the model should look like with his eyes closed. When the agent drifted, he could hear it — because the acceptance criteria lived in his head, not on the screen.

Tasks that don't fit look like this:

  • Gnarly debugging. Reading stack traces, diffing logs line by line — voice's linear, non-skimmable nature is a hard liability here. You cannot say a log comparison.
  • Design decisions that demand precision. API field naming, careful schema thought — the mouth is faster than the hand, but also sloppier. Important names and structures deserve sitting down and thinking letter by letter.
  • Production firefighting and deep water like concurrency and security. You do not want to meet an agent confidently writing a wrong migration while you are cooking. High pressure plus low-precision input plus high-stakes operations: that combination never ends well.
Diagram contrasting tasks suited to voice vibe coding, like small index-page features, against ones that are not, like complex debugging and production incidentsThe boundary of the voice workflow: the mouth expresses known intent, the eyes do acceptance; complex debugging and high-risk operations still need keyboard and screen

One line to summarize the boundary: the real bottleneck of the voice workflow was never "speaking" — it is "verifying." The smartest move in Willison's whole post was arguably the first one: get the agent to boot the preview before anything else. Mouth issues orders, eyes do acceptance — get that division right and the kitchen becomes a server room; get it wrong and it is performance art.

From typing to talking: how developer input keeps evolving

Zoom out, and this looks completely unremarkable in the history of programming — which is, at bottom, a history of steadily declining precision demands on input: assembly to high-level languages, hand-writing every line to IDE autocomplete, Copilot's line-level completion to agents shipping whole features. Each step took a little of the "express it exactly" burden off human hands and gave it to tools. Voice is the natural next point on that curve.

In the keyboard era, the bottleneck was "how fast can you type"; in the agent era, it is becoming "how clearly can you state your intent." Savor that shift: the core act of programming is moving from "entering instructions precisely" to "expressing intent clearly." Willison in his kitchen is the most literal picture of that transition — hands chopping vegetables, mouth describing requirements, code growing on the other side.

Here is a counterintuitive take: voice coding favors senior developers. Willison could do this only because he can tell at a glance whether the agent's migration is right and the templates are sane. The mouth replaces typing, not judgment. For beginners, the biggest risk of voice vibe coding is not speaking unclearly — it is not being able to read what came back. If your acceptance ability lags, the faster your mouth, the deeper the hole. Same as driving: autopilot frees your hands, not your judgment of the road.

And another: this changes what programming feels like, not what engineering is. Clarifying requirements, defining boundaries, setting acceptance criteria — those remain human jobs. You just used to do them sitting at a desk; now you can do them standing at a stove. The tool removed the "where do you sit" constraint. The "think it through" requirement did not budge an inch.

The minimal setup: five things to have before you copy him

If this made your fingers — or your mouth — itch, here is the checklist distilled from Willison's lab notes:

  • A coding-agent desktop app with a voice conversation mode. He used the Codex tab in the ChatGPT desktop app. Mind the button: "Start new voice chat," not the microphone next to the input box — two buttons, two modes, and clicking the wrong one is a completely different experience.
  • A local dev environment with one-command preview. First move: have the agent boot the dev server and open the page in a browser. This is the acceptance anchor of the whole workflow. Without it, voice is driving blind — you wouldn't even know what the agent did.
  • A small CRUD-grade feature you have already thought through. Don't challenge yourself on day one; pick something the agent is guaranteed to get right. Willison's calibration — "new model + migration + templates + import scripts" — was surgically precise.
  • A codebase you know; your own project first. You dare let your mouth replace your hands only when you know what "right" looks like. Voice coding in someone else's repo is walking a tightrope blindfolded.
  • Make peace with "busy hands, free mouth." The kitchen is optional, but "hands otherwise occupied" is exactly this workflow's sweet spot: commutes, walks, housework — dead time that can now become speak-only productive time. Programming just gained the concept of a side quest.

One last Willison-style lesson hiding in the details: he typed to get the environment ready before switching to voice. Don't let the "all voice" headline fool you — a smart workflow never makes things hard on itself. Hands when hands are due, mouth when mouth is due. The title says "almost entirely," and that almost is the honesty of a veteran.

Epilogue

When even Simon Willison starts writing code with his mouth, the question is no longer "is voice coding legit" but "does your ability to verify keep up with your mouth." The keyboard is not going anywhere — IDEs didn't kill the terminal. But "sit up straight, hands on the keyboard" as programming's default posture may be loosening.

Next time you are standing in the kitchen waiting for water to boil, consider: is there a small feature that could be spoken into existence? If so, take a page from Willison — get the agent to boot the preview first, prop up the laptop, and start cooking.

Primary source: Simon Willison's blog post "A new feature for my blog, built using my voice" (October 9, 2026) — a Newsletters index page (free weekly Substack plus paid monthly updates), built with the Codex tab's voice conversation mode in the ChatGPT desktop app against a local dev environment, running GPT-6 Astra High.

Sources

Browse projectsPublish your project

Related articles

Claude Dashboards and Motion: live dashboards and code-driven animations in beta
News
Claude Grows Two New Hands: Dashboards and Motion Enter Beta

On October 8, 2026, Anthropic put Claude Dashboards and Claude Motion into beta: dashboards built from plain-language questions on live company data, and animations generated as editable code rather than video-model footage. Docs, Slides, and Design went GA on all plans, with 45M+ artifacts created to date.

Product NewsAI CodingClaude
Illustration of a cryptographic context injection attack: Copilot CLI decrypts a malicious page and exfiltrates local secrets to an attacker
News
One Encrypted Web Page, 28 Seconds, and Your .env.prod Is Gone: Cryptographic Context Injection Hits GitHub Copilot CLI

Adversa AI disclosed CCI on October 6: malicious instructions hidden in AES-256 ciphertext trick Copilot CLI in autopilot mode into decrypting them in its own runtime and obeying the output as trusted instructions. In the demo, one encrypted page made the agent read a local .env.prod and silently exfiltrate it in 28 seconds. Microsoft's mai-code-1.1-flash fell for it half the time while GPT-5.6 models refused it outright; GitHub reproduced the chain but declined to call it a vulnerability.

Security & PrivacyAI CodingProduct News