Exa Previews ATLAS: New Benchmark Tests Search Agents on 547 Tasks Drawn from Real Search Traffic
Exa published a preview of ATLAS in October 2026, a benchmark for agentic web search with 547 tasks built from real, anonymized search demand. According to the company's own runs, no agent under one dollar per task achieves row-F1 above 0.5, and even the most expensive agents miss roughly a third of the golden answers. But all the numbers come from a company that itself sells search.
What's new
Search company Exa has published a preview of ATLAS (Agentic Tasks for Large Aggregation + Search), a benchmark designed to measure the accuracy and completeness of agent responses on real-world workflows that depend on web search. According to Exa's blog post, the benchmark combines queries grounded in real search demand with verified golden answers, plus an automated pipeline that allows both queries and answers to be refreshed as the web and models change.
The timing is no accident. Exa argues that existing benchmarks for agentic search — BrowseComp, WideSearch and DeepSearchQA — are largely saturated and partially memorized by frontier models. That leaves the field without good tools for distinguishing strong search systems from mediocre ones.
How ATLAS works
The benchmark consists of 547 deep-and-wide research tasks. The methodology, as Exa describes it:
- Real demand as the starting point: The tasks are generated from seed topics derived from clustering anonymized search demand — that is, what people actually search for.
- Entity discovery: Each task asks the system to find all entities — a person, a place or a company — that meet precise conditions.
- Multi-step enrichment: Each entity found must then be supplemented with 2–10 attributes requiring multiple search steps.
- Automatic refreshing: The golden answers are verified and can be updated automatically via a pipeline, unlike static benchmarks.
Results are reported with row-F1, the metric Exa uses in its runs. Exa has not yet published a full description of how the grader is implemented.
The numbers — reported by Exa itself
All results below come from Exa's own runs and are not independently verified:
- No agent run costing under one dollar per task achieved row-F1 above 0.5.
- Maximum-effort agents consistently beat their lower-compute counterparts.
- Even the most expensive agents missed roughly one in three golden answers. Exa interprets this as web search having significant room for improvement in use cases where completeness is critical.
- Holding the model framework fixed, Exa claims that its own search backend defines the cost-performance frontier (the Pareto frontier), with a score variation of 16 percent across different search backends.
That last point deserves its own note: Exa sells a search backend. The claim that Exa's backend in particular sits at the front of the cost-performance curve is a product claim from the interested party itself, without independent verification.
The saturated-benchmarks argument
A central part of the announcement is Exa's critique of existing evaluation tools. The company argues that frontier models now handle complex tasks far more cheaply and reliably than before, and that the bottleneck for deep-and-wide research has therefore shifted from model intelligence to the quality of the search backend — whether it actually finds relevant information in the world, Exa writes.
Existing benchmarks, the company argues, lose value over time because search providers optimize against them and the knowledge they require ends up in new frontier models. To support this, Exa cites a memorization measure defined as the share of tasks recalled by at least one of the models GPT-6 Astra, GPT-5.6 Sol or Claude Opus 5:
| Benchmark | Memorization (Exa's measurement) |
|---|---|
| ATLAS | 9% |
| BrowseComp | 48% |
| WideSearch | 59% |
| DeepSearchQA | 61% |
Exa also claims that ATLAS is cheap to evaluate: two dollars for ten full runs, versus $26–$80 for the older benchmarks and $18,000–$50,000 for Perplexity's WANDR. The last comparison should be read with care: WANDR and ATLAS use different grading methodologies, so the figures are not directly comparable.
The caveats
It is worth emphasizing how much of this story rests on a single source. All the figures — the row-F1 thresholds, the memorization percentages, the pricing measurements and the Pareto claim — come from Exa's own blog post about its own benchmark, and there is no separate technical report or independent replication yet. The company has an obvious commercial incentive to show both that agentic search is unsolved and that its own search is the best.
Several open questions remain:
- The full methodology, the grader implementation and details of the leaderboard have not yet been published — this is a preview, not a full release.
- The memorization measurement is tied to specific models (GPT-6 Astra, GPT-5.6 Sol and Claude Opus 5) and will change as the models evolve.
- The comparison with WANDR rests on different evaluation methodologies and is not directly equivalent.
Why it matters
Regardless of who wins the cost-performance race, the report points to something real for anyone building on agentic search: if Exa's runs give an accurate picture, exhaustive search at reasonable cost remains an unsolved problem. For users who need complete answer lists — market analyses, compliance checks, due diligence — it means agent results can still be significantly incomplete, and that price and quality are tightly linked.
Whether ATLAS itself becomes an industry standard depends on the methodology behind the numbers — and on whether anyone other than Exa gets the chance to test it.

