Back to Explore
NewsVibeFix 编辑部Updated Oct 1, 2026

Codex CLI's September Double Release: After Voice and Fullscreen TUI, 0.159 Teaches You to Interrupt the AI

0.156 (Sept 22) brought voice mode, fullscreen TUI and a usage dashboard; 0.159 (Sept 29) is all about control: instant_interrupt, draft recovery, tighter sandboxes. The terminal-agent race is moving from smart to controllable.

Concept illustration of a glowing dark terminal with code and an AI cursor

September belonged to OpenAI's DevDay headlines — but if you only watched the keynote, you missed the other thread: Codex CLI shipped two releases on September 22 and 29, versions 0.156 and 0.159. No keynote, no stage, just dense changelog entries. Yet those entries reveal the real battlefield of terminal agents in 2026: not who is smarter, but who is easier to live with — easier to interrupt, easier to recover, easier to keep on a leash.

This piece unpacks both releases, then talks about what the "terminal agent experience arms race" really means.

0.156 (Sept 22): voice, fullscreen TUI, usage dashboard

0.156 was the biggest feature drop for Codex CLI in months. The highlights, unpacked:

  • Voice mode on by default. F8 toggles it, /voice adjusts settings, audio runtimes bundled for Linux and Windows. Voice in a terminal agent is underrated: when you're staring at a diff thinking, speaking matches the brain's rhythm better than typing. OpenAI wouldn't default it on without internal data proving adoption — that's a risky bet otherwise.
  • Fullscreen TUI (/tui). An optional fullscreen terminal UI: transcript search, mouse text selection, right-click copy, six new themes, Mermaid diagrams and math notation in responses. Every terminal-agent veteran knows: once a transcript gets long, finding something from three days ago is hopeless. Search and selection are the first step to turning "conversations" into "assets."
  • /usage analytics dashboard. Account usage, token totals, plugin/skill activity. Billing transparency is the precondition for scaling usage — when agents start burning tokens for you, you need to know where.
  • Worktree sessions on by default. Spawn worktree sessions straight from the agent command center. Running several worktrees with agents working in parallel is already the 2026 standard posture; making it default was only a matter of time.
  • Sandbox isolation fixes. Closed real isolation gaps around Windows inbound connections and macOS file-handle writes. Permission boundaries for terminal agents have always been a dark corner; these were genuine holes.
  • 0.156.1 hotfix: same-day follow-up adding GPT-6 Sol and GPT-6 Luna to Codex's own model picker (distinct from their arrival in GitHub Copilot), with the rate-limit switch prompt now recommending Luna as the cheaper fallback. Model tiering lands in the CLI too.

0.159 (Sept 29): this one's about control

If 0.156 was "more features," 0.159 is "more control." The headliner is instant_interrupt: an opt-in capability letting new input steer Codex mid-response or during long code-mode calls.

Mind the wording: "steer" does not mean "safely cancel every operation," "undo filesystem changes," "verify task correctness," or "roll back a half-executed plan." The conservative reading: it improves interactive control within a turn, while Git checkpoints, permission boundaries, diff review, sandbox policy, and rollback procedures remain fully your job. Think of instant_interrupt as a steering wheel, not brakes-plus-airbags-plus-insurance.

But that's exactly what makes 0.159 valuable: it admits a truth — in long tasks, what people need most isn't a smarter AI, it's an AI you can interrupt anytime. The old pattern: you instruct, the agent runs, you wait, it finishes off-course, everything restarts. instant_interrupt moves course-correction from "after the run" to "during the run." Don't underestimate that shift: it cuts a long task's expected waste from "the whole task" to "one turn."

The rest of the release sings the same tune: cleaner start screen, more predictable warning behavior, better transcript navigation, native Mermaid rendering, thread-history pagination for app-server clients, Windows launch fixes, transcript-copy improvements, draft recovery, and tighter filesystem protections around approved commands. Draft recovery deserves its own spotlight: losing a long session to a crash is a pain every terminal-agent user knows, and recoverable drafts mean "sessions" are finally treated as worksites rather than chat logs.

Why "interrupt" is hard, and why it matters

instant_interrupt looks small and engineers hard. Model generation is streaming, tool calls are async, file changes are irreversible — injecting human intent into the middle of "happening" means aligning three clock domains: the generation clock (token stream), the execution clock (tool calls), the human clock (when you realize it's off-course). The tighter the alignment, the smoother it feels; misalign them and you get "I said stop and it kept writing files" ghost stories.

Shipping it opt-in is wise: the behavior boundaries are still being worked out, and defaulting it on would make everyone trip together. But the direction is right. Recall the arc: 2024 was about completion accuracy, 2025 about context length, 2026 about livability — interruptible, recoverable, auditable, rollback-able. The agent experience race is moving from smart to controllable. Same coin, different faces as DevDay's Security Cloud, Cursor's Rollouts, and Copilot's sandbox.

So what counts as a good interrupt? Three acceptance criteria: first, responsiveness. From intent to halt should be seconds, not "let it finish this part." Sub-second-class interruption is the only kind that matters — otherwise it's already written the wrong thing into your files by the time you speak. Second, clean state. After interrupting, the transcript needs a clear "interrupted here" marker, not a dangling half-tool-call that confuses the next turn's model. Third, continuability. Interrupting isn't flipping the table, it's changing direction — afterward you should be able to say "wrong approach, try plan X" rather than starting over. Run any vendor's "interruptible" claim through these three and roughly half the marketing falls away.

Meanwhile: rival terminals weren't idle either

Codex CLI's double release didn't happen in a vacuum. On September 23, Anthropic shipped Claude Code v2.1.281 — a pure stability release: self-healing for sessions stuck retrying "unexpected tool_use_id" errors, MCP-over-HTTP five-minute timeouts fixed, auto mode no longer retrying forever after a safety check declines, Write-call validation failures on slightly-off parameter names fixed. No new features, all "stop it from dropping the ball at the worst moment" fixes.

Over at GitHub, the Copilot CLI hardened local sandboxing the same week (/sandbox enable and friends), demanding slirp4netns/nsenter/iptables on Linux just to turn it on — asking for system dependencies for the sake of isolation is a statement of intent.

Put the three terminals together and the picture is crisp: September 2026's terminal-agent race isn't about who grows new arms first, but whose arms obey best. Claude Code fixed stability, Codex added control feel, Copilot added sandboxing — the controllability trio (stable, interruptible, isolatable) is becoming table stakes for terminal agents. When choosing one, don't just ask which models it supports. Ask three questions: can it recover from a crash? Can it be interrupted mid-course? Can its permissions be fenced? Only after three yeses, talk about smarts.

Appendix: how to read an agent product's changelog

A methodology that works for every agent tool's release notes. Sort entries into four buckets:

  • Capability (what's new): voice mode, fullscreen TUI. Decides the ceiling — glance and move on, don't chase.
  • Control (better governed): instant_interrupt, draft recovery, sandboxes. Decides the floor — upgrade promptly; these reduce the probability of bad outcomes.
  • Observability (more transparent): /usage dashboards, transcript search. Decides whether you can govern at all — the precondition for scale.
  • Fixes (holes patched): sandbox isolation gaps, MCP timeouts. The most boring and most important — they reveal whether a vendor is paying down tech debt or pretending it doesn't exist.

A healthy agent product has all four, and the control-plus-fixes share shouldn't be thin. If a changelog is all capability with no control or fixes, that's not "moving fast" — that's "debt piling up." 0.156/0.159 scores well: all four buckets filled, and it chewed through the hard problem of interruption — more respectable than supporting one more model.

One-line summary

0.156 and 0.159 didn't change what Codex can do; they changed what working with Codex feels like. From voice to fullscreen TUI, from usage dashboards to one-key interruption, the whole changelog answers one question: when AI starts doing long work for you, what you need most isn't it being smarter — it's you staying in charge at all times.

The 2026 terminal-agent race was never about bigger models. It's about whose transcript is easiest to search, whose drafts never get lost, whose interruption is most responsive, whose permissions are most honest. Those small things add up to the entire answer of "dare I hand it long work." DevDay does the grand narrative; changelogs do real life — and in real life, both devils and angels hide in details like "can it be interrupted." Next time you read release notes, don't just check which models were added. Read the control and fix entries — that's where a vendor's real engineering taste shows.

Sources

Browse projectsPublish your project

Related articles

Google developer documentation transformed into a structured API feeding an AI coding agent
News
Stop Letting Agents Code from Stale Docs: Google Turns Official Documentation into an API — One gcloud Line to Query, One Line to Install the Skill

On October 7, 2026, Google Developers launched the Developer Knowledge API ecosystem: official Google Cloud, Firebase, and Android docs as a programmatic source of truth, with a gcloud CLI surface, an official Agent Skill (one-line install), an MCP server, and multi-language client libraries. Why 'docs as APIs' uproots vibe coding's classic failure of models misremembering APIs.

AI CodingDeveloper WorkflowProduct Launch