LlamaIndex gathers ten document models behind one API with OpenDocRouter
The service converts PDFs and images to markdown via a single endpoint, with frontier models like Claude Opus 5.5 and open-source alternatives like MinerU2.5-Pro side by side — but all the numbers come from the company's own pages and its own benchmark.
LlamaIndex launched OpenDocRouter on October 7, 2026, a hosted document-to-markdown parsing platform that puts ten models — five frontier models and five open-source models — behind a single API, according to Unite.AI. Each model runs under a versioned "parsing recipe," and all are benchmarked on ParseBench for quality and cost. In practice, this means a developer can switch between Claude's parsing and an open-source model with a published per-page price by changing one parameter, without altering the integration.
The model lineup — and the spread in price and quality
At launch, the API offers five frontier models (Claude Opus 5.5, Gemini 3 Flash, Gemini 3.8 Flash, GPT-5.6 Terra and GPT-6 Luna) and five open-source models (Infinity-Parser2-Flash, MinerU2.5-Pro, TeleOCR, dots.mocr and PaddleOCR-VL-1.6), according to the company's model and benchmark page as relayed by Unite.AI.
The figures LlamaIndex has published show a wide spread:
- Claude Opus 5.5: 84.20 overall ParseBench score, with 93.53 on tables and 91.03 on "Faithfulness" — at $48.82 per 1,000 pages.
- GPT-6 Luna: 71.34 overall, at $0.80 per 1,000 pages.
- MinerU2.5-Pro: 70.05 overall, at $0.86 per 1,000 pages.
- TeleOCR: 57.25 overall, at $2.70 per 1,000 pages.
The picture that emerges is that the best open-source model in this lineup scores roughly on par with GPT-6 Luna at a fraction of the top model's price — a gap of more than 60x between Claude Opus 5.5 and GPT-6 Luna in per-page price. But these are the company's own numbers, not independently verified (more on that below).
How the API works
The service exposes a POST /v1/parse endpoint that accepts PDF, PNG and JPEG files, or URLs to those formats. According to the API documentation, a document can be sent as a public HTTPS URL or an uploaded file ID — both with a limit of 50 MB or 500 pages — or as inline base64 data of up to roughly 3 MB.
Synchronous runs are supported up to 50 pages; per-page responses contain markdown, status, and the "charge," so customers can compute cost per page as they go. Typed SDKs are available for Python and TypeScript via pip and npm, with source code in the GitHub repositories run-llama/opendocrouter-py and run-llama/opendocrouter-ts. The methods mirror the endpoints: parse.create, uploads.create, credits.get and models.list.
Layout grounding across models
A distinctive feature of the offering is a grounding engine that, according to the company, works with all the models: setting layout: true returns markdown with location-tagged bounding boxes and layout elements in reading order, under a common set of 15 layout classes: title, section_header, text, list_item, table, picture, chart, formula, caption, footnote, page_header, page_footer, code, form and key_value. Pages where the layout fails keep their markdown and are not charged for layout — according to the company's product description.
What rests on the company's own numbers?
Almost everything. There is no independent source in the record: the ParseBench results, the prices and the API descriptions all come from LlamaIndex's own pages, relayed through Unite.AI's same-day coverage of the launch. ParseBench is LlamaIndex's own benchmark, and no independent evaluation of the quality scores exists.
The per-1,000-page cost figures are also estimates based on tokens per page on ParseBench documents — Unite.AI notes that real costs on any given customer's pages will likely differ. And the prices are dated October 6, 2026, while the pricing setup shows Claude Opus 5.5 and GPT-5.6 Terra with recipes dated September 25, 2026, and GPT-6 Luna dated October 5, 2026. This inconsistency in version dating cannot be resolved with the available information and should be clarified by LlamaIndex.
Future plans — and open questions
LlamaIndex states that when new models the company can offer are launched, they are run through ParseBench to calibrate prompts, costs and other settings before being added. The company also plans to improve the most popular models through better prompts and better hosting for reduced latency.
The open questions are the usual ones for company-reported benchmarks: How do the models score on documents outside ParseBench? How accurate are the cost estimates in practice? And will the "all document models behind one API" wrapper hold up over time, or will the different models' weaknesses require model-specific adjustments anyway? For now, there are no independent answers — only the company's own.

