We tested Jev as a code reranker on 100 GitHub issues. Model-planned grep + Jev was our best route, though the choice of retrieval route mattered more than the choice of reranker.
The benchmark is CodeNib Base: 25 repos across C/C++, Go, Python, Rust and TS/JS. The query is the full issue text, and each repo is checked out before the fix. We built candidate pools four ways (BM25, dense embeddings, BM25+dense hybrid, and grep queries planned by a single Sonnet 4.6 call), capped each at 100, and reranked every pool with both Jev and Qwen3-Reranker-4B.
Jev answers typed questions, so we asked it one relevance score per candidate and sorted the scores ourselves.
Code-block Recall@5:
grep → Jev: 71.4%
hybrid → Qwen 4B: 65.0%
dense → Qwen 4B: 63.4%
Moving from BM25 to grep candidates raised Recall@5 by 10 to 14 points for every reranker we tried. On the same pool, Jev and Qwen 4B were within noise of each other. Jev had the higher point estimate on the grep pool (+3.3), but the interval crosses zero.
Grep supplied 17.5 candidates on average, so Jev's scoring took 0.48 s p50 and cost $0.045 across all 100 issues. The planning call is the expensive part: 4.0 s p50 and about $1.00 per 100 issues. If that's too slow, hybrid → Jev reached 67.9% at 3.4 s and $0.26 per 100 issues.
Skipping embeddings doesn't buy latency by itself, since warm dense lookup took 18 ms. This is one run per configuration on 100 issues, so treat the intervals as exploratory.
All the numbers and caveats are in the thread.
TypeSafe's Jev returns typed scores instead of text. We tested it as a code reranker on 100 real GitHub issues.
With a model-planned grep in front, Jev hit 71.4% Recall@5, vs 63.4% for dense retrieval → Qwen3-Reranker-4B, at about the same median latency (4.6 s).