TypeSafe AI's Jev doesn't chat. It just answers.
Researchers tested jev-1.13.0 zero-shot on 37 datasets spanning classification, routing, natural language inference, reading comprehension, commonsense reasoning, content moderation, legal clause analysis and rubric scoring. The model never generates free text: it returns a choice from fixed options, a rubric score, or a calibrated probability that a statement is true. The full evaluation, 346,009 requests, cost the researchers under $10. Benchmarked against Qwen3.8-27B and Gemma-4-E4B using their raw next-token probabilities, Jev scored 95-99% accuracy on IMDB, SST-2, HellaSwag and ARC, and 86.7% on the 122-language Belebele test.
This is the pitch for so-called System One models: cheap, fast, narrow tools built for the unglamorous decisions buried inside bigger AI pipelines, like checking whether an answer is grounded, flagging a policy violation, or routing a query. Jev beat Qwen outright on 27 of the 37 datasets, and none of Qwen's nine nominal leads survived the study's statistical bootstrap check, so those aren't real wins, just noise. That leaves one dataset roughly a wash. Against Gemma, Jev swept all 37.
Calibrated probabilities and a benchmark run that costs less than a coffee are a solid pitch for a routing layer nobody sees. But beating a 27B open model on a self-selected benchmark suite is a low bar when you're the vendor picking the tests.