Back to Explore
NewsVibeFix 编辑部Updated Oct 11, 2026

Decision-Model Price War Begins: Perplexity Open-Sources 27B Decider, Halves Price in Six Days

Perplexity shipped Decider, an open-weights decision model: 27B params, Apache 2.0, V1 at $0.04/M input — cut to $0.02/M six days later in V1.1. Self-reported 85.71% vs Jev's 84.51%: the closed-vs-open script replays in a new lane.

Diagram of Perplexity's Decider decision model: three typed judgments — noul, choice, score

Six Days, Price Halved

On October 1, Perplexity released Decider V1: a 27B-parameter decision model priced at $0.04 per million input tokens, output free. On October 7, V1.1 shipped: input down to $0.02/M, output still zero. Six days, price cut in half.

Even by AI price-war standards, that's explosive. Back in the day, GPT-4's price cuts were ground out quarter by quarter, and the reasoning-model price war took a full year. The decision-model category is three weeks old (Jev launched September 15), and the second player in halved its price in six days. What does that tell you? That in this new lane, cost was never the moat — price is just the opening whistle.

But Decider's real bombshell isn't the price cut. It's four words: open weights. Apache 2.0 license, 27B parameters, fine-tuned from Qwen3.8-27B. The first open-weights player in the decision-model lane. Jev opened the category, OpenAI followed with a Decisions API, and Perplexity chose a different road: handing the model directly to developers.

Decider 三种类型化判断示意图

What Decider Does: Three Typed Judgments

First, be clear about Decider's product shape, because it's a different species from an LLM API. Decider doesn't generate text. A single request handles up to 128 questions and returns three typed judgments:

  • noul: yes/no judgments, returning a probability. "Is this comment spam?" — straight to 0.91.
  • choice: pick-one-of-many, returning per-option probabilities plus the most likely item. Five intent-routing candidates, one call, five probabilities and a top-1.
  • score: ordered scoring across candidates. RAG reranking, recommendation ranking, resume screening — all this shape.

The context window is 262,144 tokens. That number is deliberate: enough to stuff "an entire document plus dozens of judgment questions" into one call. 128 questions × long context means you can pack every judgment step of a pipeline — moderation, classification, scoring — into a single call.

Note the design philosophy: this isn't "a smarter chatbot," it's "a judgment pipeline." Inputs and outputs are structured. No prompt engineering, no "please output JSON only," no parse failures. Perplexity has taken two years of hand-rolled glue layers from agent engineers and made them a native API.

That 128-questions-per-request ceiling will actually change how pipelines are written. Content moderation used to be "call the LLM to classify, call again to score, retry on failure" — three calls, three latencies, three parses. Now you stuff 128 items into one request and get 128 probability sets back. Batching moves from "the application layer assembling batches by hand" down to "natively supported by the API," with latency and cost falling linearly. For a platform moderating tens of millions of items a day, that one design detail saves a small team's annual budget.

What Open Weights Means

Apache 2.0, 27B parameters, runnable on a single GPU (quantized). That means three things. First, you can self-host — data never leaves your network. Finance, healthcare, government — anywhere data sovereignty matters and an API-shaped Jev can't enter, Decider can. Second, you can fine-tune — calibrate judgments to your own distribution with your own business data, something a closed API will never give you. Third, you can run offline — edge devices, air-gapped networks, flaky connectivity stop being blockers.

The 27B size is a canny pick: big enough for solid judgment capability, small enough to keep deployment costs sane. Perplexity clearly did the math — a decision model doesn't need an LLM's "know everything" breadth; it needs to "choose accurately given context." That's a narrower task, and 27B is enough.

The Self-Reported Report Card: 85.71% vs 84.51%

Perplexity's benchmark numbers: 11 benchmarks, 7,210 samples, Decider 85.71%, Jev 84.51%. On the RAGTruth single benchmark it's 88.80% vs 77.27% — an 11-point gap.

The ugly caveat first: these are company-reported benchmarks, not independently reproduced. Self-selected samples, self-picked benchmarks, self-configured rival setups. Everyone in AI knows what a self-reported scorecard is worth — useful as reference, worthless as acceptance. Discount by 30% until third-party reproductions land.

But even discounted, the number is worth chewing on. RAGTruth measures hallucination detection — one of the core battlegrounds for judgment models, the "fact-checker" for RAG outputs. 88.80% vs 77.27%: if that gap is real, the open Decider is substantially more accurate than Jev at "is this generated content trustworthy" — which happens to be one of the highest-frequency judgment steps in agentic systems. Every RAG agent needs an inspector.

My judgment: don't fixate on the 1.2-point gap between 85.71 and 84.51 — that's noise level. The real information is that an open 27B model has touched the closed flagship's level on decision tasks. That suggests "judgment" may be easier for open source to match than "generation" — small output space, crisp evaluation criteria. As our companion piece argued: what's easy to measure is easy to catch.

决策模型价格战:六天降价一半示意图

The Classic Script, Replayed in a New Lane

Lay out the timeline and the familiar taste emerges:

  • September 15: Jev launches, defines the "decision model" category, closed source, $870M Series A.
  • Then: OpenAI follows with a Decisions API — the giant certifies the category.
  • October 1: Perplexity ships Decider, open weights, $0.04/M.
  • October 7: Decider V1.1, $0.02/M. Six days, price halved.

Isn't this the LLM playbook of 2023–2024? OpenAI defines the category, Anthropic and Google follow, then Meta open-sources Llama, the price war starts, the ecosystem explodes. Only this time the whole script compressed into three weeks. Jev plays OpenAI, Decider plays Llama — and OpenAI itself plays the follower this time around. An amusing role reversal.

The replay isn't coincidence; it's structural. Any model category with a small output space and crisp evaluation criteria gets matched by open source extremely fast. Open-ended LLM generation is hard to evaluate and hard to chase; a decision model's output is a probability distribution — run the benchmark, and the ranking is immediate. Perplexity dared to open-source precisely because it calculated correctly: openness wouldn't dilute its edge, it would accelerate ecosystem gravity toward it.

And halving the price in six days shows Perplexity never intended to wrestle Jev on "model capability" — it's fighting an ecosystem and price war. $0.02/M input, free output: that pricing isn't making money, it's buying adoption. When developers can get a good-enough decision model for free (self-hosted) or near-free (API), Jev's $0.042 starts looking expensive. Once a price war starts, there's no going back.

Incidentally, that's a subtle stress test for Jev's $7.5 billion valuation. TypeSafe's Series A narrative rests on "the judgment layer is high-margin infrastructure," and Perplexity took six days to prove this layer's price can be driven to the floor. Sure, Jev can cut prices and follow — but the growth story a $7.5B valuation needs and the narrative of a $0.02/M price war don't share the same temperament. After capital pays for the category, it pays for share next — and in share wars, open-source players are never pushovers.

Selection Guide: Four Options, When to Use Which

The part developers care about most. There are now four approaches to decision models. Here's a blunt selection framework — conclusions first, scenarios after:

When to Use Jev

When you need out-of-the-box maximum accuracy and the judgment is extremely value-dense. Final review in financial risk control, medical triage, fraud calls on large transactions — where one misjudgment costs far more than the API bill, the $0.042/M premium is irrelevant. Also if your team has no ML deployment capacity and doesn't want to touch GPUs, Jev's API is the most hassle-free. But remember: you're paying a convenience tax and an accuracy premium — and you're locked to a single vendor.

When to Use Decider (API or Self-Hosted)

This is the default for most teams. High-frequency, medium-to-low-risk judgment steps: intent routing, content classification, RAG reranking, recommendation scoring, A/B test traffic splitting — the higher the call volume, the wider Decider's price advantage. At millions of calls a day, the monthly difference is thousands of dollars. If your data can't leave the building, or you need to fine-tune to your business distribution, self-host the 27B weights — something Jev can't offer. Offline and edge deployment: Decider is the only option.

When to Use Cloudflare Clef

When your judgment steps live on Cloudflare's infrastructure. If you're already on Workers, R2, and AI Gateway, Clef's value isn't the model — it's the integrated "zero ops, global edge, usage billing" package. For latency-sensitive, globally distributed scenarios (real-time content moderation), 50ms at the edge vs 300ms centralized are two different experiences. Choosing Clef isn't choosing a model; it's choosing a delivery method.

When to Keep Judging with an LLM

Three cases. First, the judgment needs a reasoning trace — not "choose A or B" but "why choose A," complex decisions requiring chain-of-thought that decision models can't yet provide. Second, extremely low call frequency — dozens a day — where adding a new dependency for that volume isn't worth it. Third, your judgment criteria change every few days, and editing a prompt is faster than retraining or re-tuning — flexibility itself is value.

One-line summary: high-frequency, standard, structured judgments move to decision models; low-frequency, complex, fast-changing judgments stay on LLMs. Over the next year, most agent teams will run two or three approaches at once. Hybrid is the norm.

The Judgment Layer Is Becoming Agent Standard; LLMs Just Write

Read the two pieces together and the picture is clear: Jev defined the category, OpenAI certified it, Perplexity open-sourced it and fired the price war, Cloudflare and AWS are turning it into a utility. In three weeks, "decision models" traveled the road LLMs took three years to walk: definition, follow-up, open-sourcing, price cuts, infrastructuralization.

That means the standard agent architecture is being rewritten. It used to be "one LLM does everything": it writes and it chooses, expensive and slow. The future is division of labor: LLMs handle "writing" — generating text, writing code, summarizing; decision models handle "judging" — routing, classifying, scoring, moderating, inspecting. They communicate in structured probabilities, not in natural language that needs parsing.

My prediction (marked as judgment): within 12 months, mainstream agent frameworks will support decision models as first-class citizens — the way they all support vector databases today. "Decision model" will sit alongside "LLM" and "embedding model" as a third model configuration. And "generate text with an LLM, then parse it to judge" will become the kind of code smell that gets flagged in review, like hand-concatenated SQL.

For developers, the good news is more choice, lower cost, and no single-vendor lock-in. The bad news is learning a new craft — evaluating, calibrating, and threshold-tuning decision models is a different trade from prompt engineering. Take the probability distribution from a choice call: you have to decide whether top-1 below 0.6 routes to a human or degrades to an LLM double-check; mapping score outputs to business actions means recalibrating thresholds on your own data. Unglamorous work — but it's what separates "using a decision model" from "calling a new API." That's the norm of infrastructure evolution, though: every maturing abstraction first brings a wave of learning cost, then lasting peace of mind.

The decision-model price war has begun, and six days to halve the price is only the start. When Cloudflare and AWS publish their pricing, the number goes lower still. For users, that's the best news: "choosing well" is becoming as cheap a basic capability as "computing fast."

Primary sources: OpenRouter model page (V1 released October 1, 2026; V1.1 October 7, 2026) and Perplexity's official docs (docs.perplexity.ai). Benchmark figures are Perplexity self-reported, not independently reproduced.

Sources

Browse projectsPublish your project

Related articles

Conceptual illustration of Zhipu GLM 5.3 joining the AWS Bedrock model shelf
News
After OpenAI's Agents, AWS Puts Zhipu's GLM-5.3 on the Bedrock Shelf: Chinese and American Models on the Same Cloud

Zhipu's flagship GLM 5.3 is generally available on Amazon Bedrock: a 753B-parameter MoE with a million-token context window and a leading 84.5 on the CyberGym security benchmark. AWS demoed it driving the open-source pentest agent Strix in an authorized security test. Behind the listing sits a revenue-share deal on invocation volume — Chinese and American models sold on the same cloud shelf, with Zhipu's Hong Kong shares jumping over 7% on the news.

Model UpdatesAI CodingIndustry Trends
Illustration of 2,000 AI agents collaboratively rewriting the Prime Agent codebase from TypeScript to Rust
News
2,000 Agents Rewrote Themselves in Rust: Prime Intellect's Two-Week Dogfooding Experiment

Prime Intellect had Prime Agent orchestrate 2,000+ agents to rewrite itself from TypeScript into Rust in two weeks, burning 200B+ tokens across 10,000+ sandboxes. The real story is not the 14x speedup but the honest methodology: a root agent that writes no code, verification separated from implementation, and an open admission that passing scripted parity tests does not mean production-ready.

AI CodingOpen-source ProjectsIndustry Trends