jev vs deepseek 4.1 flash on a safety classification task avg 6-10x cheaper and 6-20x faster (median to p90, open router inconsistency) yes i could train and host my own but look they did that already
very excited for simple fast classifiers like jev esp re: evals currently bloated by LLM-as-a-judge for literally everything boy stop

Sep 17, 2026 · 3:33 PM UTC

7
1
31
2,678
jev was 100% equivalent to human labels, better than deepseek in these runs because it hit output limits / did not emit a label
1
128
Sort replies: Relevant Recent Liked
Replying to @darrenangle
Wait and it was actually more accurate also?
1
2
47
they are both capable of matching human labels for this safety task at 100%, but jev did it in one shot, deepseek missed 40% either because it didn't emit a label or hit an output limit
1
1
44
Replying to @darrenangle
did they give some kinda broad access already i'm not about to start dming people for access to test
1
1
71
yeah i just signed up
1
52
Replying to @darrenangle
is it just a new output modality?
29
Replying to @darrenangle
wish it had bigger context window. would be super cool as an orchestration or search tool for whole codebases. but its only 32k
3
70
Replying to @darrenangle
The efficiency gains compared to existing models are wild. We actually went deeper on this here: x.com/AGTPinsights/status/20…
12
Replying to @darrenangle
But having it run local would make it even faster. Faster also allows you more decisions. Sometimes 100 ms is too slow especially for robots.
19