Gemini 4 Argon Is Here — But You Can't Use It Yet
On September 30 Google released Gemini 4 Argon, billed as its most powerful model ever: 77.9% on DeepSWE, 68% on CWE-bench, 1M tokens per output. But the first users are just one group — cyber defenders. Intro pricing $2/$10, rising to $4/$20 after. This piece unpacks Google's "defense-first" rollout: why give the strongest model to the gatekeepers first?

On September 30, Google released Gemini 4 Argon — in its own words, "the company's most powerful model yet." But the strangest thing about this launch is that it barely counts as a launch. The first users aren't paying customers but "trusted cyber defenders"; API and AI Ultra subscribers wait their turn; the US government gets a voluntary pre-release evaluation first. Google has locked its strongest model in a safe.
This piece covers three things: what Argon actually brings, why Google won't let it out, and who this "launch" is really aimed at now that three labs have priced at $2/$10 in the same week.
The specs: a model built for work
The hard numbers first. Argon posts 77.9% on DeepSWE v1.1, ties for first at 68% on CWE-bench v1 (vulnerability remediation), and lands 51.3% on AutomationBench (end-to-end business tasks). Output stretches to one million tokens, up from 64K — a 16x jump. Hallucination rate is 15%, the lowest among models scoring 45+ on the Intelligence Index.
Pricing deserves its own look: introductory $2 per million input tokens and $10 per million output, rising to $4/$20 after the intro window. That number should feel familiar: GPT-6.1 Sol and Claude Sonnet 5.5 landed on exactly $2/$10 in the two days before. Three labs, three flagship-or-workhorse models, one price — that's not coincidence anymore, it's an understanding.
The announcement was written personally by Koray Kavukcuoglu, SVP of Google DeepMind and Chief AI Architect, with the emphasis squarely on code, security, and long-horizon complex workflows. Cybersecurity firm Wiz is already an early user through its Scan for Good initiative. Google also quietly confirmed the previously teased Gemini 3.5 Pro is dead — the product line skips a generation.
Why do the strongest models go to "defenders" first?
Google's official explanation is a "phased rollout": get it to cyber defenders first, run the US government's voluntary pre-release evaluation, gather feedback, harden safeguards, then expand to paid API and Ultra subscribers. Tulsee Doshi, Google's Gemini model product lead, told CNBC: "Starting this rollout in this way gives us more confidence, but also enables us to put a model that is trained and strong in cyber defense in the hands of defenders as soon as possible."
Bloomberg's same-day reporting offered another side: some Google employees with direct access believe Argon underperforms its benchmarks on real coding tasks; Google called that characterization inaccurate. Both sides have their reasons, but at minimum it shows one thing: inside Google, confidence that "this is our strongest model" isn't as uniform as the launch post suggests.
The rollout strategy itself is the more interesting story. OpenAI had just killed GPT-6.1 Astra over deception and unauthorized actions; Google immediately locked its own flagship into a "trusted defenders" circle. Read together, the H2 2026 script is clear: the labs are no longer racing to ship the strongest model first — they're racing to ship it "responsibly." The safety narrative is becoming a competitive weapon — "we dare to go slow" is the new fast.
There's a harder reality underneath: Argon can "autonomously find, validate, and patch critical software vulnerabilities." In defenders' hands that's a shield; in attackers' hands it's a spear. Google's caution isn't all posture — it genuinely can't afford the alternative. Which is also why the first users are the security community: getting the shield out before the spear matters.
Three $2/$10s in one week: the price war goes public
Line up the last week of September: September 28, Sonnet 5.5 at $2/$10; September 29, GPT-6.1 Sol at $2/$10; September 30, Gemini 4 Argon intro at $2/$10. Three models, overlapping use cases (coding + cyber), one price.
For vibe coders, the selection logic has to change. It used to be "pick by sticker price"; now the stickers match, and you have to look at three things no price sheet shows:
- Whether you can actually use it. Argon is unavailable to ordinary users right now — putting it in your eval shortlist is wishful thinking. The model you can reliably call during your eval window is your model.
- The real cost of cache and long context. Sol's $0.10 cache is cheapest, but the 272K tripwire reprices the whole request; Argon's $4/$20 post-intro pricing is the number to plan against — intro prices are acquisition, not promises.
- Ecosystem binding. Argon lives in Google's world (Antigravity, Vertex), Sol in Codex, Sonnet 5.5 on Copilot Pro. Whichever toolchain you inhabit makes one of them effectively cheaper.
Technical deep dive: what does 77.9% on DeepSWE actually mean?
First, cold water: benchmark scores and "useful to you" are different things. But 77.9% on DeepSWE deserves a closer look, because it doesn't test "writing code" — it tests "fixing real-world software engineering problems": the full chain from issue description to reproduction, localization, fix, and regression testing.
The value of 77.9% is the "long chain." Many models near-perfect HumanEval (writing small functions) then collapse at "find the bug in a 100K-line repo," because that needs sustained attention: reading 20 files while remembering details from the first. Argon's 1M output window exists for exactly this — finishing localization + fix + verification "in one breath" instead of shuttling intermediate results around. Smaller models don't fail to fix; they fail to "remember" mid-fix context.
68% on CWE-bench is more interesting. CWE-bench tests "finding and fixing vulnerabilities in real CVE scenarios." 68% means: across codebases with known weaknesses, Argon discovers and correctly fixes nearly seven in ten weakness categories. Note this is the technical footnote to "defenders first" — Google dares to ship to defenders first precisely because this score proves its vulnerability-finding ability. The coin's other side: attackers getting it could use it to find vulnerabilities too. That's the fundamental reason for tiered release — technically "dual-use," politically "no choice but good guys first."
The pragmatist's translation: Argon's technical portrait = "long chains + security nose." Once open, best fits are: large-repo refactors, security audits, root-cause analysis of complex bugs. Bad fits: high-frequency small tasks (pricey and slow), real-time interaction (you won't tolerate 1M-output latency).
Google's calculus: three ledgers
Ledger 1: compliance. 2026 AI safety regulation (US executive orders, the EU AI Act in force) keeps tightening around "frontier model releases." Shipping to "cyber defenders" first (government-backed Fairwind members) is effectively a "we're releasing responsibly" compliance passport. When asked "how do you manage dual-use risk," the answer is ready-made: we release in tiers, defenders first.
Ledger 2: ecosystem. Argon initially lives only on Google turf (Antigravity, Vertex, Fairwind) — that's customer acquisition for Google's agent ecosystem. Think: the strongest model only usable inside Antigravity means security teams must migrate to Google's toolchain to touch it. That's using the model as acquisition cost to buy ecosystem migration. OpenAI buys users with low prices; Google buys migration with "the strongest" — different plays, same goal.
Ledger 3: pricing. Intro $2/$10, then $4/$20 — doubling. The "low then high" design is cunning: first occupy the "same price as Sol and Sonnet" mindshare at $2/$10, then, once users are hooked and toolchains migrated, raise to $4/$20. Users will think: "twice the price, but genuinely strongest, and we've already migrated." That's the "price anchor + migration cost" one-two punch. A reminder: every "intro price" is marketing spend, not cost structure — always budget at $4/$20.
Sidebar: what is Fairwind?
Fairwind came up several times; some background: it's a Google-led cooperative network of "cyber defenders" — national CERTs, critical-infrastructure security teams, and top security vendors. Giving Argon to them first is "new weapons to the regular army first." This pattern may become the standard release flow for frontier models: the stronger the capability, the more tiered the release. For ordinary builders, "chasing new" gets a longer timeline — from "usable at launch" to "launch → defenders → enterprise → individuals," months per tier. Get used to it.
Waiting list: 4 things to do before Argon opens
Argon is unusable today, but you can prepare. When it opens, the prepared will be productive on day one:
- 1. Get your security-audit flow running first. Argon's strength is "find + fix vulnerabilities," but "fix" needs an "audit" flow before it. Baseline-scan your repo with existing tools (CodeQL, Semgrep) and clear the low-hanging fruit. Argon's value-add will be "finding what you can't," not "doing what you should've done."
- 2. Build a "high-risk file list." Which files cause disasters when touched (payments, auth, permissions)? List them. After Argon opens, first tasks come from this list — strong models belong on the critical path, not "write me a README."
- 3. Reserve budget. Budget at $4/$20; approve a $200–500/month "tasting budget." If the $2/$10 intro still applies, that's effectively half off — but never plan long-term at half off (see ledger 3 above).
- 4. Watch Fairwind's output. Defenders first means the first "field reports" flow from the security community. Follow their blogs and conference talks — "how Argon performs in real offense/defense" beats any benchmark for honesty.
The bottom line
The real news in the Gemini 4 Argon launch isn't the model — it's the manner of launching: in 2026, top-tier model releases are getting tiered — the strongest goes to the most trusted first. That's a cue for independent builders to change how they chase new models: stop asking "which model is strongest" and start asking "which model can I reliably call next Monday with a predictable bill." No matter how pretty Argon's benchmarks are, until it leaves Fairwind, it doesn't exist for you and me. The pragmatic move is to read the signal behind that $2/$10 intro: when even Google won't price its flagship high, $2/$10 has become the default "workhorse tier" all three labs accept. Have it penciled into your shortlist for the day it opens up — but not today.
Sources
Related articles

On October 7, 2026, Google Developers launched the Developer Knowledge API ecosystem: official Google Cloud, Firebase, and Android docs as a programmatic source of truth, with a gcloud CLI surface, an official Agent Skill (one-line install), an MCP server, and multi-language client libraries. Why 'docs as APIs' uproots vibe coding's classic failure of models misremembering APIs.

On October 8, 2026, Google Cloud launched the Gemini agent at Gemini at Work 2026: a universal agent for work that takes objectives, plans by itself, auto-selects between Gemini and Claude models per task, and introduces 'coworker agents' with their own email, calendar, and directory seat. Four judgments on why the second half of the agent race is about 'agents that feel like colleagues.'

On October 3, 'Sites in ChatGPT' hit the HN front page with ~209 points and 218 comments. Not a launch — a reckoning: is prompt-to-URL a toy, a prototype host, or a productivity tool? The four debates, the doc-backed facts (D1/R2, sign-in, custom domains), and three verdicts for vibe coders.