We tested Jev against the models we use inside Context7's parsing pipeline (Gemini Flash, DeepSeek).
5 classification tasks. Results:
- 3 ties: query relevance, duplicate detection, website suitability
- 1 win: page classification — 85% vs 56%
- 1 loss: crawl-root selection — 27% vs 93%
- 10-170x faster, 3-20x cheaper
Tagging a single page: Jev wins.
Reasoning about a whole site's structure: it doesn't.
Crawl-root selection: given a URL + site nav, pick which section to crawl.
Query relevance: is the user's question about this library or something else.
Duplicate detection: are two snippets the same example.
Website suitability: is this site technical docs worth indexing.
Page classification: does this doc page have code, API ref, or info.