We used Jev to retrieve sales data 20x faster, 10x cheaper, and 12% more accurate than GPT-5 Mini.
Every time a Rox agent answers a query, it pulls from relevant transcripts, emails, CRM notes, news, and documents.
We benchmarked two ways of retrieving them:
LLM-based reranking
Jev classification
The results speak for themselves. Jev was faster, cheaper, and more accurate.
Revenue agents require thousands of points of context to understand accounts, chart relationships, and execute on sales. This research will help us serve that context with frontier-level performance for our customers.
Sep 28, 2026 · 6:41 PM UTC
19
8
78
31,731
1/ Our methodology:
We tested across 75 production-realistic queries against a benchmark set by GPT-6 Astra on high reasoning.
Jev-Score used a four-level scale (not relevant/slightly relevant/mostly relevant/directly relevant) to grade chunks, leveraging Jev’s strengths in answering fast, classification-style questions.
2
7
667
3/ Full article by @gopalkgoel1, @marcus_kuhne, Brian Xu, Calvin Yost-Wolff, and @shriram_s:
rox.com/articles/better-chea…
8
371



















