DolphinBench is a new gold standard for testing agent memory—not with trivia questions, but by making AI finish real workplace tasks that depend on facts buried in years of chat, emails, and app data.
600 tasks, 3 personas, ~500k tokens of history each. Every task: agent must succeed with the right history, fail without. You see exactly where memory systems break down.
Baseline results: best config hits 70.67% accuracy, but isn't the cheapest (costs range $61–$1,800/run, latency 32–145s). You can't just throw money or compute at the problem—trade-offs are real.
Finally, a benchmark that cares about cost, speed, and actual memory, not just recall on hand-picked answers. Dataset and code are open:
dolphinbench.ai
Get the full analysis here:
yesnoerror.com/abs/2609.2497…
// alpha identified
//
$YNE