Skip to main content
AI search API benchmark test results, comparing performance on 900 research questions. Data analysis.

Editorial illustration for New Benchmark Tests AI Search APIs on 900 Research Questions

New Benchmark Tests AI Search APIs on 900 Research Questions

2 min read

Artificial Analysis has a new way to grade the search tools that AI agents lean on when they go looking for answers online. The firm's "Search Index" puts seven providers, Parallel, Exa, Firecrawl, You.com, Tavily, Keenable, and Brave, through the same paces using one model, GPT-5.6 Luna, inside a standardized agent setup. The only variable is which search API does the fetching.

The testing runs on Stirrup, an open-source agent framework Artificial Analysis built for this purpose, with 25 runs per task to search and retrieve web pages. Three benchmarks make up the index, each weighted equally: DeepSearchQA, with 900 multi-query research questions; a 200-item BrowseComp subset built around facts that take several browsing steps to dig up; and AA-Omniscience, 600 questions spanning six knowledge domains. A tool-free baseline, where the model answers with no search at all, gives everyone a fixed point to measure against.

The numbers that come out of this setup complicate the usual assumptions about what makes a search API good, particularly when speed, cost, and answer quality don't move together the way you'd expect.

Artificial Analysis has released the "Search Index," a benchmark that measures how well search API providers work for AI agents across quality, cost, and speed.

Why this matters

For anyone building AI agents that depend on live search, this is the first time we've seen these seven providers, Parallel, Exa, Firecrawl, You.com, Tavily, Keenable, and Brave, measured against each other with the model held constant. That matters because most vendor comparisons in this space amount to marketing pages, not controlled tests. Running everything through the same GPT-5.6 Luna setup on Stirrup isolates the one variable that actually determines outcomes for developers: which search API you plug in.

The tool-free baseline is the part worth watching closely. If a provider's lift over "no search at all" is thin on DeepSearchQA's 900 questions or the BrowseComp subset's 200 hard facts, that's a real signal about whether you're paying for retrieval or just paying for latency. Founders picking a search vendor for a product launch, and researchers benchmarking agent performance, now have a shared reference point instead of anecdote.

Whether Artificial Analysis keeps this updated as providers ship changes will decide if it becomes a standard or a one-time snapshot.

LIVE22:39OpenAI slows model development amid rising cybersecurity risks