Back to Explore
NewsVibeFix 编辑部Updated Oct 10, 2026

The Same SSRF at Google, JPMorgan, and Two Governments: MCP's Structural Security Crisis

Five unrelated security teams - Google, JPMorgan Chase, Weaviate, France's DINUM and Indonesia's Tangerang City government - each independently confirmed and fixed the same SSRF flaw in their MCP servers. Independent researcher Syed Anas Mohiuddin's October update argues the bug is structural: the protocol's design, not anyone's implementation. We break down the Protocol Pivoting attack, compare the five fixes, and give vibe coders a defense checklist ahead of his 23 October MCPCon talk.

Abstract illustration of an AI agent's tool-call chain being redirected across a trust boundary, with a server icon and warning markers

Five Confirmations, One Hypothesis: The Timeline

In May 2026, independent security researcher Syed Anas Mohiuddin published an argument: the server-side request forgery (SSRF) flaws he was finding in Model Context Protocol (MCP) servers were not isolated coding mistakes but a structural defect. He gave the attack paradigm a name: Protocol Pivoting. At the time he had one solid example and an argument.

Four months later, that hypothesis received five independent validations. Google, JPMorgan Chase, Weaviate, France's interministerial digital directorate (DINUM), and the Tangerang City government in Indonesia — five security teams from five different industries — each independently confirmed and fixed the same SSRF vulnerability. They share no code and no owner. The only thing they have in common is that they all implemented the MCP protocol.

The timeline is worth laying out item by item. On 3 July 2026, CVE-2026-14540 was reserved and published on 31 July with a CVSS score of 8.0 — that was Google's MCP Toolbox for Databases (googleapis/mcp-toolbox), with Mohiuddin credited as the finder. On 25 August, Weaviate added his name to its public Security Hall of Fame. On 4 September, the hardening PR #126 for France's official data.gouv.fr MCP server was merged. On 2 September he reported the flaw to the Tangerang City government, and on 3 September the GitHub security advisory GHSA-pw2j-pj4h-f5vg was published, rated high — one day from report to advisory. JPMorgan Chase's Responsible Disclosure team confirmed his finding (medium severity), the fix is deployed, and his name appears on JPMorgan's public recognition page. Separately, security vendor Rapid7 published CVE-2026-97228 (CVSS 2.7, Low) for a related but different issue.

This is what makes the story unusual. CVEs are filed every day, but the same mistake showing up once each at a hyperscaler, a global bank, an open-source database company, a national government, and a city government — each confirmed by that organization's own security team — is an extraordinary replication density in vulnerability research. It shows Mohiuddin's May judgment was right: the flaw is not in the implementations, but in the protocol's design.

What Is Protocol Pivoting, Exactly? In Plain Language

MCP is the de facto standard by which AI coding agents call external tools: when an agent wants to query a database, hit an API, or read and write files, it does so through "tools" exposed by MCP servers. Protocol Pivoting describes how an attacker borrows this delegation semantics to cross trust boundaries. It has two failure modes; understand both and you understand all five cases.

The first is SSRF. An MCP server exposes tools to an agent, the agent calls them with arguments, and some of those arguments are URLs, paths, or endpoints. If the server takes that value and fires an outbound request without checking where it actually resolves, then who the server's network identity talks to is decided by the agent. And agents read untrusted content and act on it — a carefully crafted instruction hidden in a webpage, a document, or a tool's returned text gets passed along as a parameter. In other words, the agent becomes a megaphone: anyone who can put text in front of it can indirectly steer the server's outbound requests.

The classic landing spot is the cloud metadata service. Mohiuddin's original May example was Microsoft's playwright-mcp: its browser_navigate tool accepted any URL the agent supplied, with no SSRF protection at all, so an agent could be steered to 169.254.169.254 — the cloud instance metadata address, where temporary credentials live. Among the five new cases, the Tangerang City government's Wazuh-MCP-Server fell into the same hole: its blueteam_check_webshell tool advertised SSRF protection, but the check only rejected literal IP addresses and never resolved hostnames. Any DNS name pointing at a private, loopback, or link-local address sailed through — cloud metadata included.

The second failure mode is subtler: unsafe handling of upstream data, most visibly in logging. The five still-unfixed US federal MCP servers fall into this class — take the Department of Veterans Affairs benefits-claims server as an example. When the upstream benefits API returns an error, the server writes the full upstream response body to the log at ERROR level, with no redaction. Those bodies can contain a veteran's name, Social Security number, date of birth, and address. Note that no attacker is required: routine validation failures are enough to trigger it, and a real deployment would produce such log entries every day.

Behind both modes sits a single assumption: developers treat data crossing the MCP boundary — parameters coming in, responses coming back — as trusted, because "it came from inside the system." In an agentic pipeline, that assumption does not hold. Mohiuddin calls that sentence the thesis of the whole work; everything else is evidence.

There is a deeper implication worth flagging. Real agent deployments chain protocols: MCP for tool calls, A2A for agent-to-agent delegation, ANP and others for discovery. An attacker can hide text shaped like an A2A task instruction inside content returned by an MCP tool; the orchestrating agent reads it and, as part of normal delegation, passes it to a subagent; the subagent trusts its orchestrator and executes the instruction. If that subagent happens to hold an MCP server with an SSRF flaw, the request now originates from inside the trust boundary — with the subagent's network position and permissions. No component in the chain was "compromised" in the classic sense; each behaved as designed, and the boundary was crossed anyway. That is where the name Protocol Pivoting comes from: what pivots is not the packet, but the trust relationship between protocols.

MCP protocol architecture diagramMCP protocol architecture diagram

Why This Is a Protocol Disease, Not an Implementation Bug

The most persuasive evidence comes precisely from the case where nobody did anything obviously wrong. JPMorgan's jpmorgan-payments/ai repository contains a documentation-search MCP server: the read_documentation tool applies a domain allowlist before fetching, while its sibling tool related() takes a caller-supplied URL and fetches it server-side with no restriction. The component was forked from an AWS open-source project. Mohiuddin diffed the two versions side by side: AWS's original never dereferenced the caller's URL at all; JPMorgan's rewrite added the fetch and left out the allowlist. Two competent teams, one routine forking step, and a hole that neither codebase had on its own.

That case demonstrates three things. First, producing the vulnerability requires no carelessness — just an ordinary engineering action. Second, no scanner catches it: none of the dependencies are vulnerable, and the call graph stops at the transport boundary. Third, finding it took a human reading both versions side by side, not a tool. Mohiuddin himself admits this: he built an open-source scanner, mcp-safeguard (published on PyPI, covering six classes — SSRF, excessive permissions, prompt-injection surfaces, information leakage, authentication gaps, and lifecycle bypass), yet pattern-matching tools miss a large share of this class, his own included.

One layer deeper, the root lies in MCP's design trade-offs. MCP makes exposing a function to an agent extremely easy, but nothing along that path prompts the developer to stop and ask: who controls this argument? Where does the returned data flow? The protocol specification itself carries no mandatory security requirements, agents trust each other by default, and delegation semantics pass trust along by default. Traditional software composition analysis and dependency scanners look for known-vulnerable packages and call graphs inside code — but the dangerous input never travels the code paths those tools model. It arrives over the transport, as a tool argument the model chose, described by a tool manifest the scanner never reads. The missing redirect policy in Google's toolbox, the dropped allowlist in JPMorgan's fork, the unredacted logs on the federal servers — all of it is code a dependency scanner would pass. The vulnerability is not in any package; it is in how the server treats input it should not have trusted.

That explains why five, and why unrelated: they did not hit the same bug, they hit the same default. The protocol handed everyone the same default configuration, and that default configuration contains no security.

Five Fixes Compared: The Same Vulnerability, Five Playbooks

The good news is that the fixes are not complicated — the hard part is realizing a fix is needed. Lining up all five reveals a spectrum from "minimum viable" to "reference implementation."

Google's PR #3448 is the reference implementation Mohiuddin points people to, and the most complete of the five: it checks the resolved address at connection time rather than leaving a gap between check and use — which closes DNS rebinding attacks (resolve to a legitimate address first, swap to an internal one at connect time); it applies IP-range allow and block lists; and it rejects an unsafe base URL at startup instead of discovering it on the first request. The cost is engineering effort: this is the most work any of the five did.

Weaviate's PR #12961 took the minimal-fix route: it constrains the Google module's apiEndpoint, region, and location parameters to Google API hosts. The logic is blunt — these are the hosts the feature actually needs, and everything else is denied. Small diff, low regression risk, well suited to hot-fixing a service already in production.

France's DINUM PR #126 is titled "harden SSRF on external APIs" and opens with "Reported by Syed Anas Mohiuddin." That a national government's official server fell into the same hole as a bank's documentation tool is a fact that carries more weight than the fix details themselves.

The Tangerang City government won on speed: reported on 2 September, fixed in commit 2bbfe12 the same day, advisory GHSA-pw2j-pj4h-f5vg published on 3 September, rated high. One day for the full "report — fix — advisory" cycle — a response tempo every maintainer could learn from. But its lesson cuts both ways: the pre-incident "protection" only rejected literal IPs, a check that is meaningless in front of DNS — validating a hostname means resolving it and then inspecting the IP. It is the most basic and most commonly skipped lesson in SSRF defense.

JPMorgan's case contributes a different lesson: the same allowlist must be applied to every tool that fetches, not just the first one written. read_documentation had the check, related() did not. That "half locked, half open" state is extremely common as features iterate, and it is exactly the kind of thing worth adding to a code-review checklist.

Taken together, a competent SSRF guard has four elements: validate the resolved address at connection time, disable automatic redirects or validate every hop, route every fetching entry point through the same allowlist, and redact upstream response bodies before they hit the logs. Google did the full set, Weaviate did the most critical one, Tangerang proved response can be fast, and JPMorgan proved that missing one entry point equals doing nothing. None of them is perfect, but the five together assemble a complete answer.

MCP agent toolchain structure diagramMCP agent toolchain structure diagram

A Minimum Defense Checklist for Vibe Coders

This story is more directly relevant to the vibe coding community than it looks. Every day, people use AI to generate their own MCP servers — wiring agents to databases, internal APIs, and automation scripts, built fast and shipped fast. Mohiuddin puts it bluntly in his disclosure: MCP servers are being published faster than anyone is reviewing them, and the same two bad assumptions keep shipping with every new server. If you are a solo developer running your own toolchain, here is the checklist in priority order, each item mapped to a real case above:

  • Inventory your tool surface first. List which tools in the MCP servers you use accept URLs, paths, or endpoint addresses as parameters. Any tool where "whatever the agent passes, the server fetches" is an SSRF candidate. This is the cheapest step with the highest return.
  • Do not trust text that comes back from tools. Webpages, documents, and API responses the agent reads may all be attacker-staged. Do not let the agent feed tool output verbatim into privileged tools (file writes, infrastructure calls, outbound requests) — put at least one validation or confirmation gate in between.
  • Self-hosted servers bind to loopback by default. Japan's Digital Agency jgrants-mcp-server was found to have no authentication at all; Mohiuddin's PR #6 requires an explicit opt-in before binding to anything other than loopback — a principle that applies to everyone. Until you know exactly what you are doing, the server listens locally and only locally.
  • Copy the four fix actions as needed. Turn off automatic redirects in the HTTP client (or validate every hop); allow/deny resolved IPs to defeat DNS rebinding; cover every fetching entry point with the domain allowlist; redact upstream response bodies in logs. Google's PR #3448 is a ready-made reference implementation.
  • Scan with a tool, but do not worship it. Scanners like mcp-safeguard test through the exposed tool surface as a black box, no source needed — good for a quick baseline. But the author himself admits pattern matching misses a lot; the critical paths still need human eyes on code — especially for forked projects, where diffing the original against your version is exactly how the JPMorgan hole was found.
  • Subscribe to security advisories for the servers you use. Sixteen GitHub security advisories credit Mohiuddin as reporter, spanning command injection, authentication gaps, session hijack, and credential leaks. If a server you depend on has no advisory history at all, that absence is itself a signal.

The deeper judgment is this: vibe coding has driven the cost of writing a server to zero, but the cost of writing a secure server has not followed. AI-generated code is good at making the feature work and bad at asking "who controls this parameter." With the protocol itself offering no guardrails, that question currently has to be asked by a human — at least for now.

23 October at MCPCon: The Second Half Isn't Over

On 23 October 2026, Mohiuddin will present this cross-vendor replication map at MCPCon North America in San Jose — in his words, the May argument, five confirmations later. Several open threads are worth following after that talk.

First, the five US federal MCP servers. On 2 September 2026, he filed five findings as private GitHub Security Advisories: the VA benefits-claims server, the CMS Blue Button server, regulations.gov, USASpending, and the CDC PLACES server — all in the "second failure mode" family of missing log redaction. As of the disclosure page, all five were still in triage, unfixed. Under responsible disclosure he published no code-level detail — which means that until patches land, the public can see that the problems exist but not where they are. The remediation progress of these five is a public window into how US government AI infrastructure responds.

Second, Japan's Digital Agency PR #6. The hardening PR opened on 7 September (explicit opt-in to bind beyond loopback, caps on attachment write size) still has not been merged. An unmerged security PR next to five unpatched federal servers shows the second half of this "structural fix" has barely begun: the person who found the bugs has proven the problem exists, but getting thousands of MCP servers worldwide to actually change is a different order of work.

Third, Mohiuddin says he is still looking. His disclosure page lists 16 GitHub security advisories crediting him as reporter, with the flaw types long past SSRF: command injection, authentication gaps, session hijack, credential leaks, and bypasses of earlier fixes. He also notes merged fixes in projects such as github-mcp-server, mongodb-mcp-server, and salesforce-mcp-server, with several more reports still in triage with maintainers. SSRF is just the easiest slice to explain; the issue list across the MCP tool surface keeps growing.

For the vibe coding ecosystem, the real significance of this episode is not five CVE-style identifiers but the paradigm shift it previews for the agent era: once AI starts calling tools on people's behalf, the trust boundary is no longer drawn at "who wrote the code" but at "who controls the data." The protocol handed everyone the same default configuration, and the default configuration contains no security — after October 2026, that sentence belongs on the desk of everyone who ships an MCP server. Whether the 23 October talk pushes the protocol toward genuinely mandatory security requirements is the biggest open question this story leaves for the ecosystem.

Sources

Browse projectsPublish your project

Related articles

Concept illustration: a laptop on a kitchen counter in voice conversation with a coding agent, code on screen, cooking pots beside it
News
Simon Willison Built a Blog Feature "Almost Entirely Using My Voice" — While Cooking Dinner

Simon Willison shipped a Newsletters index page for his blog almost entirely by voice — chatting with the Codex tab in the ChatGPT desktop app while cooking dinner in his kitchen, barely touching the keyboard. When a top-tier practitioner starts coding with his mouth, voice + agents stop being a gimmick and become real productivity. A teardown of his playbook, where this workflow breaks, and the minimal setup to copy him.

OpenAI CodexChatGPTAI Agent
Illustration of a Playwright E2E regression testing workflow: a solo developer reviewing a browser test report alongside a CI pipeline
Guide
Solo Regression Testing in the AI Era: The Minimum-Cost Playwright E2E Playbook

AI's biggest fear when editing code: fixing A breaks B. This execution-layer companion to our AI Testing Strategy shows solo developers how to build an E2E regression moat with Playwright, 5 golden paths, a data-testid convention, and GitHub Actions — one hour to set up, one hour a week to maintain, inside the free tier, with copy-ready code.

Testing & QualityDeveloper WorkflowAI Coding