Independent benchmarks to help you discover, compare, and choose APIs and agents for your use cases.

San Francisco, CA
Benchmark Insights: LegalParseQA LegalParseQA is built on redline-heavy contracts: strikes inside a number, struck table cells, whole clauses deleted with nothing in their place. Hand one to an AI agent and the pipeline is: contract PDF → parser → markdown → model → answer
1
2
2
85
Gap closed measures how much of that range a parser covers. Claude Fable 5.1: 89.1% accuracy on text-layer PDFs, 95.5% of the reachable gap closed. The contracts come from RedlineBench (CC-BY-4.0), a dataset of negotiations marked up by trained attorneys.
1
17
Benchmark Insights: Company Funding Enrichment Funding stage is one of the most-used fields in GTM enrichment, and one of the hardest to trust. A provider tells you a company is at Series B, but is that last month's round or one from three years ago?
1
3
5
102
Funding rounds are scored on two boards, never averaged. Top 3 on each, statistically tied: - Freshness (rounds from last 30 days): @firecrawl spark-2, @ExaAILabs instant search and Exa agent. - Enrichment (older rounds): @firecrawl spark-1-mini, @p0 Task API , @p0 Responses API.
1
17
Benchmark Spotlight: @seltz_ai Benchmark: Company Lookalikes Search Metric: Relevant companies per $1 at 100 The Lookalike Benchmark tests whether a company search API can take one seed company and return companies just like it.
1
4
4
189
@seltz_ai also has the highest Precision@100 on the board, and is #2 on Precision@10 at 84.4%. If you're building lookalike lists for outbound or ad audiences, @seltz_ai gets you the most right companies for your budget.
1
1
17
Benchmark Spotlight: @Tiny_Fish Benchmarks: Web Search for Coding Agents (Hard Retrieval), Multi-Turn Company Search (Multi-Hop)
2
2
6
404
On Multi-Turn Company Search, where a research agent has to find every company matching 3–4 constraints across 45 questions, @Tiny_Fish fetch lifts F1 from 26.6 to 30.2, with median cost per task going from $0.212 to $0.219.
1
31