skip to content
The Weighted Average

Agentic Engineering

The Search API Your Agent Uses Costs 3× Too Much

A new 12-product benchmark shows agent search quality clusters within 10 points while cost per 1,000 tasks varies by $125.

A hand pulls a card from a library's wooden card catalog
A hand pulls a card from a library's wooden card catalog. Photograph by Daniel Forsman

Artificial Analysis published the first provider-swap benchmark of agent search APIs on Tuesday, and its most useful number is not the winner. Holding the answer model fixed at GPT-5.6 Luna, the Artificial Analysis Search Index rates twelve search products between 64.6 and 74.8 — a 10.2-point spread — while the same model with no search tool scores 33.4. Attaching any credible retrieval layer is worth 41.4 points; picking the best one instead of the worst is worth a quarter of that. But the cost of those near-identical answers ranges from $68.08 to $193.02 per 1,000 tasks, which means the wrong default costs $124,937 per million agent tasks for slightly worse answers.

That asymmetry is the operator decision. Retrieval quality has commoditized faster than retrieval pricing, and most teams picked a search vendor when the choice was about whether search worked at all. It now belongs in the same quarterly review as model selection, where the archive has already documented that harness choice swings coding-agent results by five tasks independent of the model underneath.

Same agent, different plumbing

The benchmark’s design is what makes the numbers actionable. Artificial Analysis holds every variable constant except the search provider: the same GPT-5.6 Luna candidate model at medium reasoning effort, temperature 0.6, a 25-turn budget, ten results per query, and a text-only extractor with a fifteen-second page timeout, all documented on its Search API methodology page. The agent loop itself runs in Stirrup, the open-sourced harness Artificial Analysis built for the test, so the scaffold is inspectable rather than a black box.

The index is an equal-weighted mean of three benchmarks: F1 over the 900-row public split of Google’s DeepSearchQA dataset, accuracy on a hard 200-sample subset of OpenAI’s BrowseComp, and accuracy on 600 private held-out samples of AA-Omniscience. Contamination filtering strips known benchmark sources from both the search results and the fetched page text before the model sees them. Fatal provider errors — rate limits, 5xx responses — are retried rather than published, which flatters reliability but isolates retrieval quality, the thing being measured.

Attaching any search API beats picking the best one, by four to one

Search Index score, same GPT-5.6 Luna answer model · +41.4 points from search, 10.2 across providers

No search (baseline)BraveTavily basicKeenable realtimeKeenable proParallel turboYou.comExa fastParallel fastParallel basicFirecrawlExa autoParallel advanced02040608033.474.873.773.573.273.168.467.667.167.066.565.664.6
No searchBraveTavily basicKeenable rtKeenable proPar. turboYou.comExa fastPar. fastPar. basicFirecrawlExa autoParallel adv02040608033.474.873.773.573.273.168.467.667.167.066.565.664.6
Artificial Analysis Search Index · Aug 19, 2026

Parallel Search in advanced mode leads at 74.8, followed by Exa’s auto mode at 73.7, Firecrawl at 73.5, and Parallel’s basic and fast modes at 73.2 and 73.1. Tavily’s basic tier scores 65.6 and Brave Web Search 64.6. Read the leaderboard the way a buyer should and the ranking dissolves into two clusters roughly seven points apart, with every member of both clusters doubling the no-search baseline.

The per-benchmark detail complicates any simple ordering. Firecrawl posts the highest AA-Omniscience accuracy at 73% while ranking third overall; Parallel turbo scores 75% on BrowseComp — better than Exa auto’s 74% — yet lands eighth on the blended index because its Omniscience accuracy falls to 56%. A team whose agents do factual lookup and a team whose agents do hard multi-hop browsing should not necessarily buy the same product, and the published article behind the index makes that multi-objective framing explicit rather than hiding it behind one number.

Follow the invoice, not the leaderboard

Cost is where the spread stops being academic. Artificial Analysis reports search spend and candidate-model spend separately per 1,000 tasks. Parallel fast runs $8.41 of search plus $59.67 of model tokens, or $68.08 total. Tavily basic runs $126.44 of search plus $66.58 of model tokens, or $193.02. Subtract and the gap is $124.94 per 1,000 tasks — $124,937 per million — in favor of the product that also scores 7.5 points higher. There is no quality-for-price trade being made here; one option is simply dominated.

The mechanism is visible in Parallel’s search pricing, billed at $0.001 per query in fast mode against $0.008 for Tavily basic in the benchmark’s configuration, multiplied by how often each provider makes the model search again. Tavily basic drew 15.81 queries per task to Parallel fast’s 8.41. A pricier query that also gets asked twice as often compounds: eight-fold unit price times 1.9-fold query volume produces the fifteen-fold search-cost gap the leaderboard shows.

That ratio also reframes what the search bill actually is. Search spend is 12% of a Parallel fast task’s total cost — $8.41 of $68.08 — versus 65% of a Tavily basic task. For most teams, retrieval is a rounding error inside the model bill until a per-query price and a chatty agent conspire to make it the majority of the invoice. Worse, the effect is invisible in the model dashboard, because the extra queries also inflate context: Tavily basic tasks consumed roughly 319,900 candidate tokens against Parallel fast’s 292,100. More retrieved pages mean more tokens to read, so an expensive query tier quietly raises the model bill it was supposed to be separate from.

Now stitch in a second source. OpenAI’s published API pricing charges $10.00 per 1,000 calls for its hosted web search tool, plus search content tokens billed at model rates. At the benchmark’s measured 8.41 searches per task, that hosted tool would cost $0.0841 per task — against the $0.0084 Parallel fast actually charged for the same query volume. The convenience of the built-in tool carries a 10× premium per query before a single content token is counted, and nobody publishes that comparison because it requires one vendor’s price list and another vendor’s measured query behavior.

Latency completes the picture. Time per task runs from 16.2 seconds for Parallel fast to 57.8 seconds for Firecrawl, a 3.6× spread, and Firecrawl earns its 73.5 index score by searching 15.39 times per task and spending 42.8 of those seconds inside the search tool. Parallel advanced, the quality leader, takes 38.3 seconds. For an interactive product, the top of the leaderboard is the wrong purchase; for an overnight research pipeline, latency is free and the 1.7-point gap between Parallel advanced and Parallel fast costs only $15 per thousand tasks. The same tokens-versus-outcomes discipline that governed the end of the token binge now applies one layer down the stack.

The ways this benchmark could mislead you

The most serious limitation is generality. Three question-answering benchmarks measure how well a provider supports factual retrieval and hard web navigation. They do not measure freshness on breaking news, coverage of paywalled trade publications, geographic index depth outside English, or the structured-commerce retrieval the archive examined when Onton’s product-search benchmark earned a pilot rather than a switch. A provider that trails by four index points may lead decisively on the corpus your agents actually query.

Configuration is a second caveat, and Artificial Analysis is candid about it. Every provider ran at ten results per query with defaults except a handful of documented mode flags. Real deployments tune result counts, domain filters, recency windows, and extraction depth, and those knobs move both quality and cost more than the gap between adjacent leaderboard rows. The benchmark isolates the provider by freezing exactly the parameters an engineering team would spend a week optimizing.

Third, the harness rewards a particular agent shape. A 25-turn budget with unlimited tool calls and a text-only extractor favors providers that return dense, immediately usable snippets over those whose value emerges after deeper page fetching. Providers also differ in native payload format, and the model sees each provider’s original response body — a fair test of the product as sold, but one where prompt tuning against a specific payload could shift several points.

Fourth, these are single-vendor measurements taken across two days in mid-August 2026, with Parallel fast’s numbers refreshed on August 19. Search indexes change weekly, crawl freshness varies with news cycles, and providers ship ranking updates without changelogs. Any team acting on this should treat the published figures as a hypothesis to reproduce on its own traffic, not a durable ranking — the same standard the paper applied when OpenAI disclosed a 20% monitoring-compute overhead on its own frontier workloads without external replication.

The strongest counterargument to switching at all is switching cost. Retrieval providers sit behind prompt scaffolding, result-parsing code, caching layers, and evaluation baselines. A team that moves to save $125 per million tasks and then spends three engineer-weeks re-tuning has bought nothing at typical volumes. Below roughly 400,000 tasks per year, the savings do not clear one engineer-week at fully loaded cost; above ten million tasks, they fund a headcount.

Buy the benchmark, then run your own

The honest read of this data is that agent search has become an interchangeable commodity with non-interchangeable pricing. That is exactly the market condition in which incumbency, not quality, determines what teams pay — and exactly the condition a two-week internal evaluation resolves. The evidence that would change this verdict is a provider posting a durable lift of ten index points or more on a domain-specific corpus, which would restore quality as the primary purchase criterion. Nothing in the current data suggests that is imminent.

There is a structural lesson underneath the pricing. The benchmark shows the answer model doing most of the work and the retrieval layer supplying the ground truth it cannot invent — the same division of labor visible in research on how multi-agent coding teams coordinate through shared files rather than messages, where the channel is cheap and the content is expensive. Teams keep over-investing in the orchestration and under-inspecting the plumbing.

  • Platform teams running over one million agent search tasks a year should benchmark now. The measured gap is $124,937 per million tasks between the cheapest and priciest provider at comparable quality; a two-week swap test costs far less than one quarter of the difference.
  • Teams using a hosted web search tool for volume workloads should price the alternative. OpenAI’s $10.00 per 1,000 calls against Parallel fast’s $0.001 per query is a 10× premium at the benchmark’s 8.41 searches per task, before content tokens.
  • Interactive products should optimize time per task, not index rank. The 16.2-second to 57.8-second spread matters more to a user-facing agent than 1.7 index points, and the fastest provider is not the highest-scoring one.
  • Every team should instrument queries per task. Search cost is unit price times query volume, and volume — 8.41 to 15.81 across providers — is the variable your own prompt controls.
  • Nobody should switch on a leaderboard alone. Reproduce the comparison on your own corpus with your own extraction settings; the published methodology and the open Stirrup harness make that a days-long project rather than a quarter-long one.

The paper’s view is that retrieval has quietly become the best-value component in the agent stack and the easiest one to overpay for. Twelve products, one answer model, and a $125-per-thousand spread say the market has not yet priced that in.

Sources