Everyone has been impressed by TypeSafe AI’s Jev — I wanted to see for myself how good it is — and how useful it could be for a product like
@lightfld.
I started with text understanding, because it is the foundation of everything a decision model does: it has to understand the input, and it has to understand and follow the instructions you give it.
The short takeaway: TypeSafe's tiny "Jev" classifier is surprisingly good — roughly on the level of Claude Sonnet on classification tasks, while being a couple of orders of magnitude cheaper to run.
I looked at two kinds of work. First, general text understanding: three widely used reading-comprehension and commonsense benchmarks, where Jev performed on par with Opus. Because those may well sit in Jev's training data, I also hand-authored a fresh benchmark over a single random Wikipedia article — every wrong option drawn from the text, so nothing could be memorized. There, Jev landed alongside Sonnet and Opus (it tops CommonsenseQA and MMLU-CF, ties Opus on RACE-H, and is ~150× cheaper and ~10× faster on long-context questions).
Then I tried tasks closer to sales and customer relationships. There aren't many good public benchmarks, so I dug up older customer-service and negotiation datasets. On those, Jev is a bit better than Claude Haiku and slightly below Sonnet 5. Big caveat: none of these come from real B2B sales, so the findings are directional rather than a description of what we actually do at
@lightfld .
Overall, the results look good enough that I think small, inexpensive classification models like Jev will have real uses, including in our own product.
Link to the full write-up below.