Jev is classic low-end disruption. You don't invent Jev is you are a GPU-rich hyperlab

Sep 20, 2026 · 2:54 PM UTC

18
24
804
132,469
Sort replies: Relevant Recent Liked
Replying to @cramforce
I invented it three years ago (actually, not the half-baked versions others claimed to have done a year or two ago) but for latency reasons, not cost reasons and latency matters everywhere
3
9
1,473
Replying to @cramforce
"I fear someone in a garage who is devising something completely new."
11
1,248
Replying to @cramforce
Elaborate on your definition of low-end disruption pls - curious
1,246
Replying to @cramforce
Hyper lab can also build a much better but slightly more expensive Jev. Can’t see why not
1,856
Almost a deepseek moment
686
Replying to @cramforce
But you clone it and repackage it.
411
Replying to @cramforce
Same can be said for super-efficient models like DeepSeek. Necessity is the mother of invention.
8
618
Replying to @cramforce
Low end in price, not far off in quality. Same task, 100 papers, 24 topics: Opus 5 cost 153x more per paper and agreed with Jev on 85. The 15 it disputed were not random, one returned value flags them in advance. Code and rows:
1,000 AI papers sorted into 24 topics for $0.0585. Then Opus 5 graded the labels. @nutlope's Jev paper map went viral, but the pipeline never shipped and the eval was "still running". So I rebuilt both and opened them. The first judge run came back empty: Opus spent its whole budget thinking and answered nothing. Reasoning off, second run: it agreed with Jev on 85 of 100 papers, at 153x the cost and 1.9s against 57ms per paper. The 15 misses are not random. One number Jev already returns tells you which labels to recheck. Cheap models sort. Expensive models audit only what the cheap one flags. Repost if you classify anything at scale, because the eval rows are public and anyone can rerun them with their own judge in one command. Code in the reply.
253
Replying to @cramforce
0.44s and $0.00035 turns evals from an occasional audit into a per-step primitive.
1
675
Replying to @cramforce
How I see Jev: You're not swapping out LLMs altogether, just replacing the calls that act like a smart <if/else> with structured confidence & probabilities. Fair assessment? @typesafeai @dotpem @CompleteSkeptic
566
Replying to @cramforce
think whether the closed labs do it or not it’s coming for their lunch A frontier coding agent making 100s of tool calls each turn will be eliminated sooner than later Frontier is needed for system-2 reasoning/complex tasks and their closed harnesses kept an opening via pre/post hooks where a system 1 (ideally local not jev and a 1-3B resident coder model) can do the intern work more deterministicly than auto regressive do. Agentic coding is imo biggest use case for system1+system2 thinking. Even we do system 2 two ways - thinking/overthinking/writing-notes-from-memory and in absence of knowledge reaching out to external knowledge System 1 is quick-no-brainer-decision/action/reflex - free and ultra fast; 3 jobs behind a classifier - easy-peasy, abstain, escalate (smart if else like jev) System 2 is escalation Lossless a2a communication unlike master to subagents A local model can be continuously fine tuned on a dev’s trajectories and improve itself using a dataset compiler and runtime 80:20 ideally for most tasks where devs are accepting the follow up prompt suggestion from coding harnesses and sitting their pressing enter lol
114
Replying to @cramforce
Can you please explain better
18
Replying to @cramforce
for the people by the people
296
Replying to @cramforce
Read that as "GPU-rich hype-kebab" and you know what? Yes
79