OpenAI Decisions API: Turning an Agent's "Branch Decisions" Into a Single 150ms Call
At DevDay 2026, OpenAI launched the Decisions API: no free-text generation — just pick one answer from preset options, with probabilities. Claimed ~150ms per decision, priced on input tokens only. The twist: this "decision model" category was defined first by tiny startup TypeSafe AI's Jev, about two weeks before OpenAI. For vibe coders, the real win is turning agent loops' most annoying chore — branch decisions — into one format-drift-free API call.

What happened
At DevDay 2026 on September 29, OpenAI launched a new API: Decisions. It's currently in limited preview, with the company saying access will expand within days; a decisions guide page went live in the developer docs on October 1.
Sam Altman demoed it live in the keynote: a computer-use agent whose every "what next" judgment ran through the Decisions API. On October 6, Simon Willison wrote an llm-openai-decisions plugin based on the official docs — the first round of independent testing, so regular developers can now verify the company's claims hands-on.
The takeaway up front: this isn't another "smarter chat model." It's a purpose-built interface that judges and doesn't chit-chat. Once you internalize that positioning, every detail below falls into place.
The docs guide makes the positioning explicit: Decisions isn't a "lite version" of a chat model but a standalone product line — the input and output schemas are designed for judgment from the ground up. It won't chat with you and it won't explain its reasoning (at least not through this interface). It answers exactly two things: which option, and with what probability.
What it actually is
In one sentence: it focuses GPT-6 Luna's intelligence on "answering a question you define."
It generates no free text. You hand it a question plus a set of preset answers; it picks exactly one of the presets and returns it with per-option probabilities. Input supports text and images.
The docs define three question types: yes/no (binary judgment, e.g. "is this user message asking for a refund?"), choices (pick one, e.g. "which queue should this ticket be routed to: support / billing / sales?"), and scores (rating, e.g. "rate this comment's toxicity 1–5"). The stated uses are equally blunt: content classification, request routing, choosing an agent's next action.
Pricing: $0.10 per million input tokens — input only, output free. Speed: the company claims roughly 150ms per decision, about 10x a standard Luna call. Note the caveat: "roughly 150ms" and "10x" are the vendor's numbers; independent large-scale measurements haven't caught up yet. File that under "to be verified" — more at the end.
Image input deserves a special mention because it's the clearest differentiator over Jev's text-only approach. Visual decisions — "does this uploaded photo contain a receipt?", "which of these four damage categories does this car photo show?" — previously required a full vision-model call with all its latency and cost. A 150ms vision judgment call collapses an entire category of moderation, triage, and inspection workflows into something you can run inline, per item, without batching.
OpenAI didn't invent this category
The interesting backstory: the "decision model" category wasn't pioneered by OpenAI but by a small startup called TypeSafe AI. Its product, Jev, launched around September 15 — roughly two weeks before OpenAI.
Jev's specs: $0.042 per million input tokens, free output, 70–500ms latency, text-only input, per-option probabilities, 64K-token context. Compare OpenAI's version: input pricing more than double Jev's ($0.10 vs $0.042), but with image input and the GPT-6 Luna base model behind it.
A giant following a category defined by a tiny startup is worth savoring. It says two things. First, the "decision" need is real — not a pseudo-demand dreamed up by some big lab: a small company validated PMF first, and the big lab followed. Second, AI infrastructure is now innovating fast enough that "the right to define a category" is no longer monopolized by big tech. A team of a few dozen people defines an API category; two weeks later OpenAI announces its own — a script unimaginable in the cloud-computing era.
For indie developers, that's good news: your niche isn't far from "defining the next API category." The Jev team is probably about your size.
The price math is worth doing carefully. Jev: $0.042 per million input tokens, free output. OpenAI: $0.10 per million input tokens, output equally free (because there is no output). The premium buys two things: image input, and GPT-6 Luna's base-model intelligence. For pure text classification, Jev is usable today and cheaper; for judgments that need vision — say, deciding which category a user-uploaded image belongs in — OpenAI is currently the only option. Don't let keynote glow distort the decision; choose on input modality and latency requirements.
There's also a strategic read on the timing. OpenAI validating a startup's category two weeks after launch is the fastest "fast follow" in AI infrastructure history. It tells you the moat in this layer isn't the idea — it's distribution and base-model quality. For builders, that means: prototype on whoever is cheapest today, architect so you can swap tomorrow. The interface shape (question plus options plus probabilities) is already converging into a de facto standard, which is exactly what makes swapping painless.
What it means for vibe coders: control flow becomes one API call
Now look back at the most annoying part of building agents. Write an agent loop and you'll find dozens of branch decisions inside it: is the user asking to cancel the order? Should this error be retried? Which tool gets called next? Today's approach: burn one full LLM call per decision — write the prompt, wait for generation, parse the output, handle format drift, retry on drift.
Every experienced agent developer has done this janitorial work: telling the model to "output JSON only," regex-extracting fields from the reply, retrying when the output isn't valid JSON. An agent's overall reliability often dies exactly here, in format drift at branch points. You spend three days on prompt engineering only to discover that 3% of production failures come from the model whimsically adding an extra newline one day.
Decisions compresses the whole janitorial routine into one API call: send the question plus options, get back a structured choice with probabilities. No free text means no format drift; ~150ms latency means branch decisions stop being the slow bottleneck in agent loops; input-only pricing means the cost math is trivial.
That's cost and reliability dropping together: cost falls on per-token pricing with no output charges; reliability rises on the design decision to lock the output space shut. A model that can only pick from five options cannot hand you back an essay — and "cannot hand back an essay" is exactly the guarantee you've wanted on all those late debugging nights.
One level deeper: this is LLM APIs "speciating." For years, every task went through the same chat interface — poetry and classification shared one hammer. Now the hammer is splitting: generation belongs to generative models, decisions belong to decision models. Behind the interface split, "intelligence" is being decomposed into priced, measurable standard parts. For builders, standard parts mean composable, replaceable, comparable — the mark of a maturing ecosystem, and the market structure indie developers love most: nobody gets to lock you in with a single black-box interface.
One easily overlooked feature: probability outputs. Decisions returns per-option probabilities, which means you can set thresholds — high confidence executes automatically, low confidence escalates to a human or to a full model. That's a graceful-degradation engineering pattern: previously you wrote the sampling, voting, and threshold logic yourself; now the interface hands you probabilities and branch policy becomes configuration instead of code. Agent behavior shifts from "hardcoded if-else" to "tunable policy" — a genuine step up in maintainability.
Think about what this does to agent architecture over time. Today's agents interleave "thinking" and "deciding" in one model call, which is why prompts balloon into multi-page instruction manuals. Once decisions are cheap, fast, and structured, the natural refactor is a two-tier loop: a decider that routes at 150ms a pop, and a thinker invoked only when judgment needs nuance. The agent gets faster not because the model got smarter, but because the dumb parts stopped pretending to be smart.
Don't refactor just yet: open questions
First, the 150ms and "10x" figures are OpenAI's own claims. Simon Willison's plugin test just landed; independent large-scale latency data doesn't exist yet. Whether your scenario sees 150ms or 500ms, you'll have to measure yourself. Between keynote numbers and P99 latency there is always the real world.
Second, limited preview. Working today doesn't mean your key is on the allowlist next week. Hold production migration plans until GA — or at least until the promised "expansion within days" materializes. Breaking changes are routine at the preview stage.
Third, price. $0.10 per million input tokens is more than double Jev's. If your workload is pure text classification, Jev ($0.042, free output) is usable today and cheaper. OpenAI's edge is image input and the Luna base — choose on merit, not on the logo.
Fourth, and most important: locking the output space buys you an expressiveness ceiling. Judgments that need nuanced reasoning — "is this code review comment questioning the architecture or the implementation?" — will probably still need a full model. Decisions replaces "branch decisions," not "thinking." Learning to distinguish which of your judgments are classification and which are reasoning is the most important lesson before adopting it.
The pragmatic move: start with the single most painful branch decision in your agent — usually routing or classification. Run it for a week and compare: how much did latency drop? Did format-drift retries disappear? Did the token bill move? The numbers will tell you whether this is genuinely useful or just another keynote demo. Good tools aren't afraid of measurement; the ones afraid of it aren't good tools.
And keep one eye on the category, not just the vendor. Jev proved the demand, OpenAI validated the shape, and the interface is converging. Where there's a converging interface, there will be a third, fourth, and fifth implementation — quite possibly open-weight ones you can self-host. The winners in this layer won't be whoever launched first; they'll be whoever makes the decision call cheapest, fastest, and most boring. Boring infrastructure is the best infrastructure.
Sources
Related articles

On October 7, 2026, Google Developers launched the Developer Knowledge API ecosystem: official Google Cloud, Firebase, and Android docs as a programmatic source of truth, with a gcloud CLI surface, an official Agent Skill (one-line install), an MCP server, and multi-language client libraries. Why 'docs as APIs' uproots vibe coding's classic failure of models misremembering APIs.

On October 8, 2026, Google Cloud launched the Gemini agent at Gemini at Work 2026: a universal agent for work that takes objectives, plans by itself, auto-selects between Gemini and Claude models per task, and introduces 'coworker agents' with their own email, calendar, and directory seat. Four judgments on why the second half of the agent race is about 'agents that feel like colleagues.'

On October 7, 2026, GitHub announced via Changelog: starting with CLI 1.0.94-0, the /model command discovers models in your local Ollama instance, listed alongside configured and cloud models. Discovery doesn't auto-enroll — each model needs manual confirmation — and models must support tool calling and streaming. GitHub also teased intelligent routing, and clarified: a local model neither enables offline mode nor disables telemetry.