Codex Issues a "28-Day Pledge": One Improvement a Day, or a Usage Reset for Everyone
On October 4, OpenAI's Codex engineering lead Tibo made a public pledge: for 28 days, ship one clear improvement every day for most Codex and ChatGPT Work users — on days the team fails, everyone gets a usage reset. Day 1 brought ~50% faster GPT-6 Astra / 6.1 Sol by default, Day 2 made Auto-review free for ChatGPT sign-ins, Day 3 put GPT-6 in the chat tab. This is an experiment in turning product iteration into a daily series.

A "military order": 28 days, one improvement per day
On October 4, Thibault Sottiaux — Tibo, OpenAI's Codex engineering lead — made an unusual public pledge on X: for the next 28 days, ship one clear improvement every day for most Codex or ChatGPT Work users; on any day the team fails, everyone gets a full usage reset. Note the clever design — the "penalty" isn't an apology letter, it's free quota handed to users. The pressure sits entirely on the team's own shoulders.
On October 5, the OpenAI Developer Community opened an official tracking thread with daily updates. As of October 7, the shipped list reads: Day 1, GPT-6 Astra and GPT-6.1 Sol ~50% faster by default on subscriptions; Day 2.1, Auto-review free for ChatGPT-account sign-ins; Day 2.2, simplified API rate-limit qualification; Day 2.3, a Meetings plugin for notetaking; Day 2.4, Decisions API in public beta; Day 2.5, rate-limit resets for everyone; Day 3, GPT-6 in the ChatGPT chat tab with a brand-new Intelligent UI.
Why does this deserve its own story? Because it isn't a routine product update — it's an experiment in turning "iteration speed" itself into a marketing event. The AI coding-tool war has spread from "whose model scores higher" to "whose shipping rhythm feels more addictive."
Day 1: GPT-6 Astra / 6.1 Sol ~50% faster by default
Day one of the pledge delivered pure performance: GPT-6 Astra and GPT-6.1 Sol roughly 50% faster by default through subscriptions. Per the announcement, it's an inference-side optimization — nothing to configure, rolling out within two hours.
The coverage is the interesting part: not just Codex and ChatGPT, but every partner product supporting Sign in with ChatGPT — officially named: OpenCode, Pi, Amp, and Devin. Sign in with your ChatGPT account in those third-party tools and you get the speedup too. The practical takeaway for vibe coders: model-speed dividends are spreading from OpenAI's first-party apps across the whole "ChatGPT sign-in ecosystem," and "supports ChatGPT sign-in" is quietly becoming a selection criterion when picking tools.
The community already has measured data. One user benchmarked before and after with DIGBENCH: GPT-6.1 Sol on task P-1 went from 34m36s / 45.2M tokens to 9m40s / 6.4M tokens — time and tokens both down, suggesting shorter reasoning paths, not just faster decoding. A single sample, but it points the same direction as the official claim.
Day 2: Auto-review goes free — a second pair of eyes at no cost
Day two shipped the security headline: Codex's Auto-review is now free for ChatGPT-account sign-ins and no longer counts against plan usage. Enable it under Settings → General → Permissions → Auto-review; CLI users can set approval_policy = "on-request" with approvals_reviewer = "auto_review" in config.toml.
A quick explainer: when Codex needs to do something "outside the sandbox boundary" — a blocked network request, files outside its folder, tools with external effects — it used to stop and wait for your manual approval. Auto-review sends a second agent to review those requests for you: it weighs your original request, the environment, and the action's likely impact, screening four risk classes — private data to untrusted destinations, credential/token probing, persistent security weakening, and irreversible destructive operations. Actions already allowed inside the sandbox get no extra check, so it doesn't nag.
The feature shipped back in late April, but reviews consumed plan quota — "hire a bodyguard, pay his salary yourself" — so many users turned it off. Making it free turns "secure by default" from slogan into default configuration.
The numbers come from the alignment team's April 30 evaluation: internal March data showed approval interruptions ~200x fewer than manual-approval mode, with 99.1% of boundary-crossing requests approved; synthetic tests showed 99.3% of prompt-injection attempts blocked and 90.3% of overeager actions (the agent going beyond user intent while trying to help) caught, plus 96.1% on MonitoringBench (Hard). One illustrative snapshot of 10,000 internal actions: 9,280 never left the sandbox and needed no review; 720 went to the reviewer, 713 approved, 7 denied — of which 4 found a safer route on their own and only 3 ended up asking a human.
OpenAI added its own caveat: Auto-review is not a security guarantee. Red-team tests found the reviewer can be misled, and it can't stop a model deliberately hiding its intent. It adds a check at the approval step; it doesn't replace human oversight. The direct advice for vibe coders: turn it on now that it's free, but don't treat it as a free pass — before letting an agent near your database, your backups and permission config are still the last wall.
Day 3 and beyond: GPT-6 enters the chat box, on a 40M-user base
The October 7 Day 3 drop took the campaign to ChatGPT itself: GPT-6 landed in the Chat tab alongside Intelligent UI — conversations can now return charts, forms, tappable controls and other interactive elements instead of just text. That's a key step in the chat box's evolution into a workbench, sharing the same model base as Codex.
The same day, Tibo posted another number on X: Codex and ChatGPT Work together hit 40 million active users. The metric needs unpacking: in June the company said Codex passed 5 million weekly users; in an August 25 interview Tibo said ChatGPT Work reached 20 million users. The 40M is a combined figure with no stated measurement window — it evidences the scale of the agent user base OpenAI is assembling, not either product's standalone growth rate. Even so, 40 million is itself the confidence behind the 28-day pledge: the bigger the base, the lower the marginal cost of each daily ship, and the longer the word-of-mouth lever.
The community's two faces: measurers vs. skeptics
The tracking thread's comments are the best window into how this experiment is landing. Roughly two camps.
The measurers are seriously verifying each day's delivery: before/after benchmark posts, confirmations that partner products got the speedup, readings of the "performance improvement" as an inference-side win. Their logic: whatever the marketing gloss, the thing is real and the experience genuinely got better.
The skeptics are watching the substance: Day 2 packed five "2.x" mini-items (2.1 through 2.5), read by some as padding the count; some call the thread a rebroadcast of subscription-product promo tweets; one developer used the moment to demand fixes instead — "rather than daily UI tweaks, fix the Codex VS Code task-queue bug the community already patched four times." Their logic holds too: once "daily shipping" becomes the KPI, teams favor what's easy to ship over what's hardest and most valuable.
Both camps point at the same question: what happens after day 28? If this is a one-off marketing sprint, everything reverts when the heat fades; if it hardens into the team's shipping rhythm, that's a genuine organizational upgrade. Too early to call — but day 28's delivery quality will say more than day 1's.
An overlooked detail: the reviewer model changed too
Behind free Auto-review sits another cost equation. According to AI media AICatchup's roundup, OpenAI's developer account said on X that auto-review in the ChatGPT app and Codex CLI is being upgraded from GPT-5.4 to GPT-5.6 Luna — and combined with Luna's new pricing, review cost is expected to drop roughly 10x. Caveat: the "10x" currently exists only as an X announcement with no independently re-checkable primary source, so take it as directional.
But the direction itself is telling: high-frequency, low-per-call-value work like safety review is being deliberately migrated to cheaper models. It's the same logic as the Zed 1.22 story we covered — strong models for hard tasks, cheap models for grunt work (review included). OpenAI built this "model tiering" into its own product: the main agent runs GPT-6 Astra/Sol, the reviewer runs Luna. Vibe coders can copy the division of labor directly: the agent running security scans in your CI doesn't need the same expensive model as the one writing code.
One fine-print note from the official docs: free auto-review applies to ChatGPT-account sign-ins; direct API usage or enterprise workspaces with their own policies may not be covered. Free doesn't mean unconditional — check your plan terms before assuming.
Our take: agent tools enter the "operational density" era
Plotting the "28-day pledge" on the H2 2026 map, one shift is clear: the AI coding-tool battleground is moving from "model capability" to "operational density" — who ships more often, whose feedback loop is shorter, who makes users feel "this tool gets better every day." Cursor built that feeling with rapid releases; OpenAI just pushed it to the extreme with a public pledge under full community supervision.
Three direct implications for vibe coders. First, the moment to bet on the Codex ecosystem is now: 50% faster, free review, partner-wide coverage — the subscription's value is visibly rising through these 28 days. Second, turn on Auto-review now that it's free, but remember its boundary — it reduces interruptions, it doesn't replace backups. Third, watch the "Sign in with ChatGPT" niche: OpenAI is turning the subscription into a cross-product passport, and whether your third-party agent tools support that sign-in directly affects whether you taste dividends like these.
One observation slot left open: on day 28, will Tibo present a handsome report card, or owe everyone a "missed homework" reset? Either way, users win — which is precisely the pledge's cleverest design: it tied "the team ships" and "users benefit" into the same event. If you build products, this mechanism is worth copying.
Sources
Related articles

On October 8, 2026, Google Cloud launched the Gemini agent at Gemini at Work 2026: a universal agent for work that takes objectives, plans by itself, auto-selects between Gemini and Claude models per task, and introduces 'coworker agents' with their own email, calendar, and directory seat. Four judgments on why the second half of the agent race is about 'agents that feel like colleagues.'

On October 7, 2026, Google Developers launched the Developer Knowledge API ecosystem: official Google Cloud, Firebase, and Android docs as a programmatic source of truth, with a gcloud CLI surface, an official Agent Skill (one-line install), an MCP server, and multi-language client libraries. Why 'docs as APIs' uproots vibe coding's classic failure of models misremembering APIs.

On September 30, 2026, Bitdefender launched AI Guardian in public beta: a security layer for autonomous AI agents that verdicts every tool call, file access, and credential use as allowed, flagged, or blocked. First on macOS, free during beta, supporting Claude Code and OpenClaw. Why this 'agent behavior firewall' arrives right on time for vibe coders.