Vercel AI Gateway Ships Stealth Reasoning Model Glyph Cluster, Free During Stealth: A Hands-On Testing Guide and the Data-Privacy Red Line
Vercel added stealth reasoning model Glyph Cluster to AI Gateway (Oct 7): coding plus long-context analysis, function calling, streaming. Pro/Enterprise teams with purchased credits use it free during stealth (stealth/glyph-cluster), selectable in Claude Code, Codex, Cursor. Why indie developers should benchmark it, the gateway's failover and budget value, and the limits: text-only input, no structured outputs, no ZDR — prompts may train the model, so set a data boundary first.

Vercel has quietly slipped a new model into AI Gateway: Glyph Cluster, free during its stealth period, available as stealth/glyph-cluster. The official changelog entry went live on October 7, worded with restraint — but behind the restrained wording sits a genuinely friendly signal for independent developers: a reasoning model tuned for coding and long-context analysis is now available to test at zero cost.
This article unpacks the confirmed facts from the changelog: who gets it free, what the model is actually good at, how to plug it in, and — most importantly — the data-privacy boundary that didn't make it into the headline.
The facts first: who gets it free, and how
The free tier comes with a gate. Only teams on Pro and Enterprise plans that have purchased AI Gateway credits can use this model for free during stealth. Note the wording: it's the stealth-period usage that's free, not AI Gateway itself — you need to already be a paying-plan customer who has bought credits. What pricing looks like after stealth ends isn't stated in the changelog.
This is classic Vercel playbook: let real users run the model on production traffic first, collect feedback, and refine pricing. From the user's side, it amounts to an officially sanctioned "new reasoning model trial channel" — you get to learn the model's temperament on real work without spending extra.
Three integration paths are offered: the AI SDK, OpenAI-compatible Chat Completions / Responses APIs, or simply selecting the model inside coding agents like Claude Code, Codex, and Cursor. For anyone already using those tools, the switching cost is near zero: pick it from a dropdown and your next turn is Glyph Cluster doing the work. That "change a model name and try it" design is precisely the ideal setup for head-to-head comparison — same task, same prompt, different model, and the differences speak for themselves.
One CLI command to get going
The official setup is refreshingly short:
npm i -g vercel@latest
vercel ai-gateway setup
Two commands and AI Gateway auth and configuration are in place. No elaborate onboarding, no application form to fill out. For anyone wanting to validate quickly, the gap between reading the changelog and firing the first request can be compressed to minutes.
What this model is actually for
The changelog positions Glyph Cluster with unusual specificity: a reasoning model built for coding and long-context analysis. The keyword is "reasoning" — this isn't a rapid-fire chat model, but one tuned for multi-step work: planning, synthesizing information, quantitative reasoning, and comparing across large inputs.
Translated into daily developer life: it can review code, explain failures, debug, suggest modifications, and reason through implementation approaches. Notice the ordering of that capability list — review and "explaining failures" come first. That suggests the training objective weights reading and understanding existing code, and pinpointing root causes, heavily — not just generating code from scratch.
That distinction is worth savoring. Most coding models on the market are competing on "generation": prompt in, code out, faster and faster. But in real development, time spent reading code dwarfs time spent writing it — onboarding onto an unfamiliar repo, chasing a bizarre production bug, reviewing a teammate's PR. Those are the black holes that devour developer hours. A reasoning model that's stronger at "understanding and explaining" hits exactly the most expensive part of the work.
Function calling: reasoning wired into your own tools
Glyph Cluster supports function calling, plus streaming output. That means agents can connect its reasoning to their own tools and systems: querying databases, calling internal APIs, reading and writing files. Reasoning stops being an isolated block of text and becomes a chain that can drive action.
For anyone building agent applications, this combination (reasoning model + function calling + streaming output) is the standard answer. Streaming keeps the UX snappy, function calling lets the model actually do things, and reasoning capability makes sure it thinks before it acts. All three are indispensable — and the changelog names all three.
Why indie developers should take this update seriously
What solo teams and indie developers lack has never been models. It's a coding reasoning model that's both cheap and capable.
Let's do the blunt math: top-tier reasoning models doing deep code analysis burn tokens at multiples of chat scenarios. An indie developer doing code review, architecture deliberation, or running multi-step agent tasks can easily blow past expectations on daily reasoning-token spend. Expensive models bill monthly, but the anxiety accrues daily — hesitating before every "send" over whether a question is "worth the premium model" is itself a productivity tax.
That's where the free stealth period matters: it zeroes out the cost of experimentation. You can go all-in on two categories of tasks: long-context analysis — throw in an entire module's code and ask for review and refactoring suggestions; and multi-step reasoning — root-causing a complex bug, or reasoning through an implementation plan across files. The free window exists to answer one question: on your real workflow, can this model replace the one you're currently paying for?
My advice is concrete: run it on three tasks you genuinely handled last week. Not benchmark questions, not toy examples — your own code, your own bugs. Benchmark scores are for other people; your own task pass rate is your answer. If it performs comparably to your current primary model on your tasks, the only remaining variable is the post-stealth pricing.
The AI Gateway layer is undervalued
The headline star of this news is Glyph Cluster, but the thing truly worth noticing may be AI Gateway itself.
What AI Gateway does is simple: one entry point to many models. For indie developers, that solves three real pains. First, failover: when a model goes down, gets rate-limited, or slows to a crawl, requests automatically route to a fallback — your production service doesn't go down with it. Second, budget control: the credits mechanism gives you a clear cap and visibility over AI spend, instead of an inscrutable bill at month's end. Third, comparison convenience: same code, same interface, swap a model name and you're running A/B tests — no per-vendor integration layer to write.
The counterintuitive part: the more models there are, the more valuable a unified gateway becomes, not less. Many people think "I only use one or two models, why do I need a gateway." But the model landscape shifts fast — this month's best coding model may be dethroned next month. Without a gateway layer, every model switch is a mini-refactor; with one, switching models is editing a string. The arrival of stealth models like Glyph Cluster proves the point: it landed inside the Gateway on day one, and Gateway users get zero-cost trials. That's infrastructure compounding.
But first, read the three limitations carefully
The changelog states its limitations with disarming honesty — honest enough to deserve a line-by-line reading.
First, text-only input. No image or file inputs. You can't hand it a screenshot of an error or a UI walkthrough directly — everything has to be converted into text first. For pure code workflows that's fine, but if your habit is "screenshot + one question" debugging, this model can't take that.
Second, no structured outputs. You can't demand responses in a fixed JSON schema. That's a genuine constraint for agent workflows: steps requiring strictly structured output (like parsing results into fixed-format task cards) need you to wrap your own parsing and validation around it, or leave those tasks to models that support structured outputs.
Third, and most important: no ZDR (zero data retention). The changelog's plain meaning: your prompts and responses may be used for training and model improvement. Free during stealth has a price, and the price is your data.
On data privacy, let's be blunter
What does no ZDR mean? It means every line of code, every error stack trace, every architecture description you send to Glyph Cluster could theoretically end up in future training data. For personal side projects that's usually fine; for commercial code, there's a line that must be drawn in advance:
- Never feed code containing secrets, tokens, or real user data. That's an iron rule regardless of ZDR — but the consequences are worse without it.
- Desensitize proprietary algorithms and core business logic before testing. Abstract the key logic into an equivalent but de-contextualized version to validate the model's capability. Validating capability and using real code are two different things.
- Set a team policy before the team starts using it. A free stealth model makes it all too easy for teammates to casually paste production code in. Before the free window opens, get explicit in the team about what may be fed in and what may not.
This isn't alarmism. Nearly every historical case of "casually fed code to an AI" leakage started with "it didn't seem like a problem at the time." A free stealth model plus a frictionless workflow is precisely the combination most likely to lull people into dropping their guard.
A pattern worth thinking about: stealth models are multiplying
Glyph Cluster isn't the first model to launch in stealth. Over the past year, the "ship anonymously, collect real feedback, reveal identity later" release pattern has become increasingly common. For model vendors, stealth launches yield the most honest evaluation data — users who don't know which model it is carry no brand bias; for users, the stealth window often comes with free or discounted access, a low-cost window into frontier models.
But there's an implicit exchange in this pattern: you're helping vendors tune models with real data, and vendors are buying your feedback with free credits. The trade itself is fair — the problem is that many people don't realize they're sitting at the table. Knowing the rules before you sit down is a different game from sitting down blind.
What to do now
If you're a Vercel Pro / Enterprise user who has purchased AI Gateway credits, the action list is short:
- Run
npm i -g vercel@latestandvercel ai-gateway setupto wire up the gateway. - Switch the model to
stealth/glyph-clusterin Claude Code, Codex, or Cursor — or call it directly through the OpenAI-compatible API. - Run three tasks you genuinely handled last week through it: one code review, one bug root-cause analysis, one implementation-plan deliberation. Record how it differs from your primary model.
- Set the data boundary inside your team first: which code may be fed in, which may not. Policy before free rein.
If you're not in the target group (say, on the Hobby plan), this news still matters as a trend signal: stealth models plus unified gateway access are becoming the standard debut format for new models. The next time a similar free window opens, you'll want your workflow already in "swap a model name and try it" shape — and the cost of getting there is building the gateway layer now.
One final judgment: whether Glyph Cluster is worth using long-term is too early to call; but whether the free stealth window is worth an afternoon of hands-on testing — the answer is yes. The upside of free is bounded by your time; the downside is zero cost for first-hand feel. In a 2026 where models iterate on a weekly cadence, that feel is itself an asset — it decides whether, at the next model reshuffle, you're switching proactively or waiting passively.
Primary source: Vercel Changelog: Glyph Cluster is now available in stealth for free on AI Gateway (published 2026-10-07).
Sources
Related articles

On October 3, 'Sites in ChatGPT' hit the HN front page with ~209 points and 218 comments. Not a launch — a reckoning: is prompt-to-URL a toy, a prototype host, or a productivity tool? The four debates, the doc-backed facts (D1/R2, sign-in, custom domains), and three verdicts for vibe coders.

On October 7, 2026, Google Developers launched the Developer Knowledge API ecosystem: official Google Cloud, Firebase, and Android docs as a programmatic source of truth, with a gcloud CLI surface, an official Agent Skill (one-line install), an MCP server, and multi-language client libraries. Why 'docs as APIs' uproots vibe coding's classic failure of models misremembering APIs.

One prompt, six unsupervised hours: GPT-6 Astra finished a three.js visualization of Calvino's 55 Invisible Cities in 53 minutes for $10, while Claude Opus 5.5 took 85 minutes, 6 parallel subagents and $74 — and declared a 'jump' for design tasks. A field report on vibe coding's limit case, model selection, and the token ledger.