Back to Explore
NewsVibeFix 编辑部Updated Oct 8, 2026

An "OpenRouter" for Document Parsing: LlamaIndex Launches OpenDocRouter with 10 Models Behind One API

On October 7, LlamaIndex launched OpenDocRouter: one POST /v1/parse endpoint backed by 10 document-parsing models — five frontier closed-source models plus five open-source ones, switchable at will. Its companion ParseBench shows Claude Opus 5.5 scoring 84.20 at $48.82 per thousand pages, while GPT-6 Luna scores 71.34 at just $0.80 — a 60x price gap for a 14-point difference. Document parsing finally has a public, transparent quality-vs-cost menu.

Hand-drawn illustration of document parsing: a magnifying glass scans a document covered in OCR recognition boxes, with arrows pointing to structured table output

On October 7, 2026, LlamaIndex announced a new product on LinkedIn: OpenDocRouter. The name might suggest yet another parser upgrade, but what it actually does is different — it puts 10 document-parsing models behind a single POST /v1/parse endpoint, letting developers switch freely between all 10 with one API key.

The playbook should look familiar: what OpenRouter did for large language models, OpenDocRouter is now doing for document parsing. Over the past year, the biggest pain point in RAG applications has not been retrieval algorithms but step zero — turning PDFs, scans, and financial tables into clean text. That step is dirty and expensive, and every parsing vendor has a different API, so switching providers means rewriting your integration code. With this launch, LlamaIndex is effectively declaring that parsing models can be routed like interchangeable commodities, just like LLMs.

Not a parser — an OpenRouter for parsing

First, the product shape. OpenDocRouter is not an eleventh parsing model; it is a unified routing layer. You send a request to POST /v1/parse, specify which model to use in the parameters, and any of the 10 models does the actual work behind the scenes.

The 10-model lineup deserves a close read. Five frontier models: Claude Opus 5.5, Gemini 3 Flash, Gemini 3.8 Flash, GPT-5.6 Terra, and GPT-6 Luna. Five open-source models: Infinity-Parser2-Flash, MinerU2.5-Pro, TeleOCR, dots.mocr, and PaddleOCR-VL-1.6. The curation is deliberate — the most expensive closed flagships on one end, the community's best-regarded open-source parsers on the other, with lightweight Flash-class models in between. From "money is no object" to "every cent counts," both ends of the spectrum are covered.

It attacks a very real switching cost. Previously, comparing two parsing services meant separate signups, separate SDKs, and separate response formats — and most people just stayed with whichever one they tried first out of inertia. With a unified interface, comparison becomes a parameter change, and the cost of A/B testing drops toward zero. That is especially friendly to indie developers and small teams: you can finally evaluate like a big company — run everything, look at quality and the bill, then decide.

A 60x price gap for 14 points: ParseBench puts cost negotiation on the table

Launching alongside the product is ParseBench, a parsing benchmark built on roughly 2,000 pages of human-verified real enterprise documents with 169,000 verification rules. The "enterprise" qualifier matters — it does not test clean academic papers but the real messy stuff: financial reports, contracts, and invoices with chaotic layouts and nested tables.

Once the scores landed, the interesting part was not who came first but the price-to-score table. Claude Opus 5.5 scored 84.20 overall and 93.53 on tables — a decisive lead — but costs $48.82 per thousand pages. GPT-6 Luna scored 71.34 at just $0.80 per thousand pages. The open-source MinerU2.5-Pro scored 70.05 at $0.86 per thousand pages. Do the math: a 60x price difference buys a 14-point score gap.

That number deserves a pause from everyone building RAG. Does your use case actually need those 14 points? If you are extracting key figures from financial statements, 93.53 versus low-70s on tables could be night and day, and the $48 is money well spent. But if you are stuffing 100,000 pages of internal wiki into a vector store for semantic search, the downstream difference between 70 and 84 may be unmeasurable — while the bill differs by 60x. OpenDocRouter's real value is not routing itself but the fact that it puts the quality-versus-cost menu out in the open for the first time: no more guessing from sales pitches and blog benchmarks; the vendor pinned price and score together.

A pragmatic tiered strategy writes itself: bulk, low-sensitivity documents go to cheap models like MinerU2.5-Pro or GPT-6 Luna; table-dense, number-sensitive financial reports get routed to Claude Opus 5.5 alone. Route by document type instead of picking one model for everything — that is how a routing layer is meant to be used.

Billing and privacy: per-token pricing, frontier models at cost, failures free

Several billing details show LlamaIndex courting developers. First, per-token billing instead of crude per-page bundles — you pay in proportion to the content you actually get back. Second, frontier models at cost with no markup: whatever Claude or GPT calls cost is what you pay; LlamaIndex takes no cut in the middle. Third, a 5% top-up fee. Fourth, failed, cache-hit, and blank pages are all free.

The last point is easy to overlook, yet it is the most practical saving in parsing workloads. Real-world document collections are full of blank pages, duplicate uploads, and failed-then-retried parses — pure waste under per-page pricing. Free failures plus per-page independent retry billing drive the cost of retrying to zero: you can finally retry a failed page three times without doing mental accounting on the bill.

The privacy commitment is equally direct: documents are not retained by default. Parse caches live only 24 hours, are encrypted at rest, and can be deleted anytime. For teams handling contracts, medical records, or financial data, "not retained by default" beats any compliance badge — data that never lands on disk has no leak surface. And the 24-hour cache means re-parsing the same document can hit the cache for free, the same logic as free failures: saving money for real workflows.

A homegrown grounding engine: layout:true gives every span of text coordinates

Beyond routing, OpenDocRouter ships a homegrown grounding engine: add layout: true to the request and every text span in the response carries a bounding box, plus document structure described by 17 layout labels — headings, tables, headers and footers, lists, and image captions, each in its place.

Why does this matter? Because the citation experience in RAG lives or dies on grounding. "The answer is on page 5" versus "the answer is in this highlighted box on page 5" are two different levels of trust. The former still makes you hunt; the latter jumps straight to the highlight. Layout labels pay off most on tables and nested structures: when the parser knows this is a table and that is its header row, downstream structured extraction stops guessing.

Here is the judgment call: anyone can copy a routing layer — it is just a unified interface plus a model list. But building bounding-box-level grounding into parse output and aligning it across 10 models' formats is OpenDocRouter's real moat. Models will be replaced and prices will shift; a stable, high-precision layout-understanding layer is the long-term asset.

Dividing labor with LlamaParse: one for speed, one for precision

The first question many asked: what happens to LlamaParse? LlamaIndex's answer draws a clean line: OpenDocRouter handles fast onboarding and free routing; LlamaParse keeps doing enterprise-grade tuning and schema extraction.

Translated: validating an idea fast, comparing models, avoiding lock-in — use OpenDocRouter. Production systems that need fine-tuning on specific layouts and strict schema extraction from parse output (invoice fields, contract clauses) — that is LlamaParse's job. This is not the left hand fighting the right; it is splitting the experimentation layer from the production layer into two products. For LlamaIndex, OpenDocRouter is the top of the acquisition funnel — get developers in through routing, and heavy users will naturally settle into LlamaParse.

The engineering constraints came with the announcement too: 50MB and 500 pages per request; synchronous responses under 50 pages, asynchronous above; per-page independent billing and retries. Those numbers signal it was designed for production workloads from day one, not as a demo toy.

Our take: the dirtiest, most expensive job in RAG finally has a public menu

Back to the opening claim: in the whole RAG pipeline, document parsing is the dirtiest, most expensive, and most underestimated step. Dirty because layouts are endlessly weird; expensive because good models charge per page without blinking; underestimated because everyone obsesses over retrieval algorithms while forgetting that information lost in step zero never comes back.

OpenDocRouter's significance is not technical but structural. It turns parsing from a heavy decision — pick a vendor, sign a yearly contract — into a light one: change a parameter, pay per page. Now that "60x price gap for 14 points" is pinned publicly on a benchmark, it gets much harder for any parsing vendor to profit from information asymmetry. That is pure upside for buyers.

Three pieces of practical advice for developers. First, write your parsing layer behind a routing interface now — do not bake any single vendor's SDK into your business code. Whether or not you use OpenDocRouter, replaceability is an asset in itself. Second, benchmark on your own documents before going live; do not trust generic scores — your layouts are your ground truth. Third, build the habit of routing by document type: cheap models handle the bulk, expensive models handle only the 5% that deserves them. Cost optimization for parsing will become as standard a RAG engineering practice as LLM routing already is.

Unite.AI's October 7 coverage independently confirmed the core facts of the launch. For LlamaIndex, the move is unsurprising — from LlamaParse to OpenDocRouter, it is evolving from "a parsing tool" into "the infrastructure layer for document intelligence." The second half of the RAG arms race may not be about whose model is bigger, but whose pipeline is cheaper, more transparent, and easier to swap.

Sources

Browse projectsPublish your project

Related articles

A laptop screen showing a website signup page inside a browser
News
ChatGPT Sites Hits HN's Front Page: Prompt-to-Website — Toy or Productivity?

On October 3, 'Sites in ChatGPT' hit the HN front page with ~209 points and 218 comments. Not a launch — a reckoning: is prompt-to-URL a toy, a prototype host, or a productivity tool? The four debates, the doc-backed facts (D1/R2, sign-in, custom domains), and three verdicts for vibe coders.

AI CodingProduct LaunchIndie Development
Google developer documentation transformed into a structured API feeding an AI coding agent
News
Stop Letting Agents Code from Stale Docs: Google Turns Official Documentation into an API — One gcloud Line to Query, One Line to Install the Skill

On October 7, 2026, Google Developers launched the Developer Knowledge API ecosystem: official Google Cloud, Firebase, and Android docs as a programmatic source of truth, with a gcloud CLI surface, an official Agent Skill (one-line install), an MCP server, and multi-language client libraries. Why 'docs as APIs' uproots vibe coding's classic failure of models misremembering APIs.

AI CodingDeveloper WorkflowProduct Launch