AI Moves In to Guard the Grid: Anthropic Launches Critical Infrastructure Defense Program CIDP
On October 8 Anthropic launched the Cyber Mission: the Critical Infrastructure Defense Program (CIDP) — deploying frontier Claude models, engineers, and threat research to power grids, water, and transport via 11 founding partners — alongside a free OSS scanner. With Project Glasswing's 129,000 verified vulnerabilities as evidence that AI has collapsed the cost of finding bugs while fixing stays slow, CIDP marks the turn from AI writing code to AI guarding code.

On October 8, Anthropic launched a long-term program called the Cyber Mission, aimed at protecting critical infrastructure and open-source software from cyberattacks. The program runs on two tracks: the Critical Infrastructure Defense Program (CIDP), and a free AI vulnerability-scanning service for open-source projects. This article focuses on the former — because it may be the first time a lab has deployed frontier coding models, as an organized effort, into operational technology (OT) settings like power grids, water systems, and transportation networks, rather than keeping them inside a chat window.
A note on sourcing first: Anthropic's official blog was unavailable during this reporting round, so the facts below rest on three in-depth third-party reports — the-decoder's October 9 coverage, intelligibberish's October 9 daily roundup, and a long-form breakdown on dev.to. Sources are cited throughout so readers can trace everything back.
An asymmetric ledger
To understand CIDP, start with the hardest sentence in Anthropic's announcement: "The cost of exploiting vulnerabilities has dropped, while verifying, disclosing, and fixing them is slow." That is not PR gloss; it is the premise the entire program is built on.
Over the past two years, AI's impact on the offense-defense balance has been visibly lopsided: finding vulnerabilities keeps getting cheaper. Models can sweep through more code in hours than a human auditor could review in weeks, and writing working proof-of-concept exploits keeps getting easier. On the other end — confirming whether a bug is real, judging how bad it is, writing the patch, and actually getting the patch deployed — almost nothing has sped up. The dev.to deep dive puts the bottleneck bluntly: the bottleneck is no longer discovery. It is triage, prioritization, and patching, and all three still depend on people.
Critical infrastructure magnifies this asymmetry to the extreme. Power grids, water supplies, and transport networks run on operational technology built to last decades. That equipment often cannot be taken offline to patch, so known vulnerabilities can sit unresolved for years. Worse, these systems are maintained inside a closed ecosystem: the equipment is proprietary, changes are risky, and one wrong move can take down a substation or a plant. Securing a substation and securing a web app are fundamentally two different crafts.
Open-source software tells a parallel story: nearly all modern software depends on open-source code, while many critical components are maintained by small volunteer teams with no security budget. Attackers have industrialized. Defenders have not. That is the structural problem Anthropic is going after.
Which makes its forecast worth pausing on. Anthropic offers its own prediction: "Our forecast is that in two years, AI will favor defense… But in the near term, that may not be true." The first half is a roadmap; the second half is honesty: right now, AI made offense cheap first.
What CIDP actually does
CIDP works like this: Anthropic gives operators of critical infrastructure access to frontier Claude models, on-site engineers, and threat analysis. But note the delivery choice — it does not go directly to utilities or water companies. Instead, it works through 11 founding partners, putting these capabilities in the hands of the "trusted providers" operators already rely on.
The 11 are: Accenture, Booz Allen, CrowdStrike, Deloitte, Dragos, Hitachi, Insane Cyber, Nozomi Networks, Palo Alto Networks, PwC, and Rockwell Automation. The composition of the list is more informative than the list itself: Dragos, Nozomi Networks, Rockwell Automation, and Hitachi are OT specialists who speak substations and factory floors; Palo Alto Networks and CrowdStrike are the big security vendors; Accenture, Booz Allen, Deloitte, and PwC are consultancies that translate technical fixes into processes enterprises can execute; Insane Cyber is a smaller, more specialized player.
Why take the indirect route? The dev.to analysis nails it: Anthropic knows it has neither the credentials nor the relationships to secure a substation. The OT world's rule is "who is accountable when something breaks," and a model lab knocking on the door would get ignored by a utility. But if the capability arrives via people already inside the customer's plant — Dragos or Rockwell Automation — the trust chain works. This is a classic borrow-the-boat strategy: Anthropic supplies models and engineers; partners supply relationships, domain knowledge, and accountability.
The playbook isn't new, either. CIDP builds on a program launched in June aimed at US state, local, tribal, and territorial governments, which has reportedly since reached more than half of US states. The Cyber Mission effectively extends that model from government administrative networks into the physical world's infrastructure.
From "writing code" to "guarding code"
Zoom out, and CIDP looks like a landmark turn in the history of AI programming. The last three years told a clean three-act story. Act one: AI writes code — Copilots put completion inside the editor. Act two: AI reviews code — code review, vulnerability scanning, and auto-fix suggestions entered the CI pipeline. Act three, now: AI guards code — models are no longer confined to static files in a repo but are deployed into settings where failure has physical consequences, taking part in ongoing defense.
What makes CIDP special is that it is the first program to send frontier coding models into OT settings as an organized deployment. Earlier AI security tooling mostly stayed on the IT side: scanning dependencies, reviewing PRs. The OT bar is far higher — protocols are proprietary, equipment can't be rebooted, and a false positive can mean downtime. So Anthropic didn't ship a generic scanner to the market; it chose the heavy model: embed engineers, do threat research, deliver through partners. That tells you the lab itself has realized that at the OT layer, "here's an API" isn't enough. Defense is a service, not an endpoint.
For vibe coding practitioners, the turn carries a direct implication: the code you generate with AI today may be audited by another AI tomorrow. Code "production" and code "review" are being taken over by the same class of models, and CIDP means the endpoint of "review" now extends into production systems themselves. The people writing code and the people guarding it may soon use the same base models — with opposite objective functions. That is a signal every AI coding tool builder should take seriously: the quality of your tool's output will eventually be scored by defense-side models.
129,000 vulnerabilities: the cost structure is being rewritten
The other number disclosed with the Cyber Mission is the key evidence for how seriously to take this round. Between April and July this year, Anthropic's earlier vulnerability-hunting effort, Project Glasswing, "uncovered at least 129,000 verified software vulnerabilities," with "more than 33,000 rated critical or high severity." Third-party reporting adds that Anthropic believes the true impact to be at least five times higher.
Don't read that number as a scorecard yet. What 129,000 really says is: the marginal cost of vulnerability discovery has collapsed. One project, four months, over a hundred thousand verified bugs — an output density unimaginable five years ago. It is the plainest evidence for the earlier claim that AI is rewriting the cost structure of offense and defense, cheapening the "find the bug" side first.
But Anthropic also admitted the awkward part: Glasswing did not achieve a sufficient reduction in overall cyber risk. Finding bugs got easy; fixing them did not. That is exactly why the Cyber Mission exists in this shape — piling on more discovery capacity wouldn't move the risk curve. So the center of gravity shifted from "find" to "fix" and "defend": CIDP puts capability in the hands of partners who can actually drive remediation, instead of shipping another vulnerability list.
Folded in alongside Glasswing is the Cyber Verification Program, now formally expanded into three tiers. Defense Access covers SOC and incident response, malware reverse engineering, and vulnerability analysis; Red Team Access covers authorized penetration testing; Specialized Access covers flight operations, power grids, telecom, interbank transfers, and government administrative networks. Available models include Claude Opus 5.5, Claude Sonnet 5.5, Claude Mythos 5.1, and future releases. Note the design logic: vetted defensive security professionals can apply for access to more capable models with reduced blocking classifiers — Anthropic is opening a carve-out in its safety policy for defenders, letting models do more in defensive contexts. That is a subtle but important policy turn: the model's "safety restrictions" are starting to distinguish whether the user is an attacker or a defender.
The honest parts, and the parts left unsaid
The most worthwhile passage in the dev.to breakdown is its assessment of Anthropic's candor. This announcement does contain some unusually frank admissions: that Glasswing did not reduce risk enough; that the free scanner's reports are model-generated, ship without human review, and will contain errors; that AI cannot fix the parts of critical-infrastructure defense that are about physics, access, and organizational capacity. That level of candor makes the remaining claims easier to evaluate — a lab willing to say its last project didn't fully succeed before pitching the new one earns a different kind of credibility.
But several things are left unsaid. First, money and scale: how many embedded engineers, how many operators covered, what the bar for partners is — none of it disclosed. Eleven founding partners sounds like a lot, but the US has thousands of utilities; initial real-world coverage will necessarily be small. Second, measurement: Anthropic defines success as water, power, transport, and communications keeping running under attack from capable adversaries, with fewer exploitable paths and faster recovery — a goal measured in years, unverifiable in the short term. Third, geography: the disclosed partners and government programs are all US-centric, and the announcement doesn't say whether this "critical infrastructure defense" is a global agenda or an American one.
There is a deeper question, too: as defensive models get stronger, won't the offensive side benefit in lockstep? Anthropic's answer is the "two years" forecast — the logic being that once AI accelerates fixing and prevention, defense's scale effects will outrun offense. But that prediction rests on an assumption: that defensive organizational capacity (the processes that allow downtime for patching, the people who can read the reports) can keep up with model capability. And Glasswing's lesson was precisely that organizational capacity is the slowest link. The two-year timeline reads less like a forecast than a bet waiting to be tested.
Three tiers of access: defenders are being "treated differently"
The Cyber Verification Program's three-tier design deserves its own close look, because it reveals Anthropic's new answer to "who should get model capability." Defense Access serves SOC and incident response, malware reverse engineering, and vulnerability analysis — the daily work of the defending side: shifts, forensics, sample teardown. Red Team Access serves authorized penetration testing — the "legal offense": people paid to find holes in their own systems. Specialized Access is the most sensitive tier: flight operations, power grids, telecom, interbank transfers, government administrative networks — settings where failure propagates directly into the physical world or the financial system.
All three tiers share a model roster: Claude Opus 5.5, Claude Sonnet 5.5, Claude Mythos 5.1, and future releases. The point isn't the models themselves but the reduced blocking classifiers. Translated: questions a safety policy would block for an ordinary user can be permitted for a vetted defender. Anthropic is starting to tier its safety policy by identity — who you are and what you do determines how much capability the model opens to you.
The turn is debatable. Supporters will say it was inevitable: locking the strongest analytical capability in a safe does nothing to attackers — they were already routing around restrictions every which way; the people actually delayed are the defenders who need models to help them tear apart malware and reproduce vulnerabilities. Critics will worry the vetting mechanism itself becomes the bottleneck: who gets to certify you as a "defender"? Are the criteria transparent? Does it become a big-institution privilege? Public information hasn't answered these yet, but the direction is clear: model capability allocation is moving from "equal for all" to "tiered by use."
For developers outside the US, there's a practical sense of distance: CVP's and CIDP's disclosed footprint is US-context all the way down — the partner list, the state government programs, the infrastructure types under protection all read American. That doesn't weaken the reference value, but it means that if you work in critical infrastructure or security outside the US, what you can directly use in the short term is the thinking more than the channel: the three-tier design, the "deliver through trusted providers" path, and the "admit fixing is harder than finding" prioritization are all directly copyable.
The other track, in one sentence
The Cyber Mission's other track is OSS Scanner, a free AI vulnerability-scanning service for open-source projects that periodically scans enrolled projects and suggests fixes. (This site already covered it in detail in its October 8 reporting; not repeated here.)
Worth adding is the funding side: Anthropic says it has funded the Python Software Foundation, Alpha-Omega, the OpenSSF through the Linux Foundation, and the Apache Software Foundation, and launched the Defender Advantage Fund (0xDAF) in August to back pilot programs and keep the scanner free. Maintainers can also apply for free Claude Max subscriptions through Claude for Open Source. It is a combined play: model capability for partners, funding and tooling for the open-source ecosystem.
Why this matters to the vibe coding community
Back to the three-act framing. AI writing code solved the "production" problem; AI reviewing code is solving the "quality" problem; CIDP is attempting the "consequence" problem — when AI-generated code runs inside power grids and water plants, who makes sure nothing goes wrong? The program pushes the responsibility boundary of frontier models from "help you write" to "help you guard." For this community's builders, that means three things.
First, security is moving from "after-the-fact audit" to "infrastructure." If the CIDP model works, any serious production system may come with an AI defense layer by default, the way it comes with CI today. The tools that write code and the tools that guard it will converge on the same base.
Second, OT settings are the next hard but high-value frontier for AI programming. Proprietary protocols, machines that can't be rebooted, costly false positives — these constraints are exactly what general coding models handle worst. Whoever solves "patch safely on machines that can't restart" gets the ticket into critical infrastructure. CIDP's choice of heavy-touch delivery says the moat here isn't the model itself but domain knowledge and delivery capability.
Third, defenders are getting a policy dividend. The Cyber Verification Program's three tiers are, in essence, directionally opening stronger model capability to the defensive side. That is a signal: future model capability may no longer be "same rights for everyone" but tiered by purpose. People doing security and infrastructure get the better tools first.
Finally, back to Anthropic's "two years." Whether you read it as roadmap or marketing depends on whether programs like CIDP can prove one thing over the next two years: that AI can not only find vulnerabilities faster, but make them disappear faster. Until then, the most honest description remains the second half of that sentence — in the near term, offense is still cheaper. And admitting that is precisely the weightiest part of this round.
Sourcing note: Anthropic's official blog was unavailable during this reporting round. Facts rest on cross-verified portions of three third-party reports — the-decoder (Oct 9), intelligibberish (Oct 9 daily roundup), and a dev.to deep dive; direct quotes are from Anthropic's official announcement as relayed by those reports. Figures and the partner list follow the original reporting.
Comments (0)
Sources
Related articles

In an October 9 release-notes announcement, Google ended new sales of Gemini Code Assist Standard and Enterprise subscriptions: existing subscriptions auto-renew through 2026, with auto-renewal ending in 2027. From June's individual-tier shutdown to October's full sales halt, Google spent five months absorbing AI coding into the Antigravity enterprise agent platform. The era of standalone AI coding tools is ending; indie developers must now choose ecosystems, not single products.

Codex CLI 0.161.0 ships in-session MCP login, new Bedrock capabilities, and explicit Daybreak opt-in; GPT-5.5 retires from Codex on all plans October 14. This piece breaks down the four confirmed changes and gives a pre-October-14 migration checklist — from grepping your configs to pinning models.

Microsoft's homegrown coding model MAI-Code-1.1-Flash now runs on-device: 3-bit quantization, full 256K context, zero inference charges for local calls — but Microsoft recommends 120GB+ RAM. Third-party measurements put the quantized build at 70.80% on SWE-Bench Verified versus 72.6% full precision. What "zero inference fees" really means for indie developers, and why Copilot's router is the actual moat.