Every team evaluating Linkup asks the same first question: how do you perform against LLM native search and the competition?
Fair question. On public benchmarks, we lead: SimpleQA and SealQA-0, state of the art.
But raw accuracy only tells you so much. That's why we launched verticalized benchmarks that map to real use cases: GTM (People, Company), Legal (RedlineBench), Finance (FinSearchComp).
Each links to an open repo with the dataset, queries, and methodology. Re-run it yourself.
But what really matters is what works for you. Picking a provider on a public leaderboard alone is how you end up with failures in production: a benchmark measures general performance, not yours.
So we go one step further and help you build a custom eval on your own queries and data, guide included.
Public benchmarks show who leads in general. A custom eval shows who leads for you.