[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-typesafes-jev-skips-chat-and-just-picks-an-answer":10,"sections":46},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":35,"tags":36,"sources":41,"feedback":45,"feedback_at":22,"cost_usd":45,"total_tokens":45},8646,"typesafes-jev-skips-chat-and-just-picks-an-answer","TypeSafe's Jev Skips Chat and Just Picks an Answer","A commercial model that never generates text outperforms Qwen on most of 37 benchmarks by returning structured decisions instead of prose.","TypeSafe AI's Jev doesn't chat. It just answers.\n\nResearchers tested jev-1.13.0 zero-shot on 37 datasets spanning classification, routing, natural language inference, reading comprehension, commonsense reasoning, content moderation, legal clause analysis and rubric scoring. The model never generates free text: it returns a choice from fixed options, a rubric score, or a calibrated probability that a statement is true. The full evaluation, 346,009 requests, cost the researchers under $10. Benchmarked against Qwen3.8-27B and Gemma-4-E4B using their raw next-token probabilities, Jev scored 95-99% accuracy on IMDB, SST-2, HellaSwag and ARC, and 86.7% on the 122-language Belebele test.\n\nThis is the pitch for so-called System One models: cheap, fast, narrow tools built for the unglamorous decisions buried inside bigger AI pipelines, like checking whether an answer is grounded, flagging a policy violation, or routing a query. Jev beat Qwen outright on 27 of the 37 datasets, and none of Qwen's nine nominal leads survived the study's statistical bootstrap check, so those aren't real wins, just noise. That leaves one dataset roughly a wash. Against Gemma, Jev swept all 37.\n\nCalibrated probabilities and a benchmark run that costs less than a coffee are a solid pitch for a routing layer nobody sees. But beating a 27B open model on a self-selected benchmark suite is a low bar when you're the vendor picking the tests.","[\"ai models\",\"benchmarking\",\"llm evaluation\",\"typesafe ai\"]","2026-09-30T04:00:00.000Z","2026-09-30T16:57:42.013Z","2026-09-30T16:57:47.656Z","published",null,[24,30],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"Attribute the benchmark to its actual source (arXiv:2609.37647v1) with a named publication\u002Fdate instead of leaving it unsourced, and drop or verify the 'independent researchers' framing since the source doesn't establish the authors are unaffiliated with TypeSafe AI, the vendor of the product being evaluated.","resolved",{"id":31,"reviewer":32,"round":33,"reason":34,"status":29},"publisher-r2","publisher",2,"The win counts don't add up: Jev is said to beat Qwen on 27 of 37 datasets, but Qwen is then credited with 'nine wins' — 27 + 9 = 36, not 37, an unexplained internal inconsistency in the reported results.","ai",[37,38,39,40],"ai models","benchmarking","llm evaluation","typesafe ai",[42],{"name":43,"url":44},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.37647",0,{"sections":47},[48,51,55,59,64,69,73,78,83,87,92,97,102,107],{"name":49,"slug":35,"count":50,"latest_published_at":18},"AI",5180,{"name":52,"slug":53,"count":54,"latest_published_at":18},"Security","security",791,{"name":56,"slug":57,"count":58,"latest_published_at":18},"Policy","policy",417,{"name":60,"slug":61,"count":62,"latest_published_at":63},"Deals","deals",284,"2026-09-29T21:00:00.000Z",{"name":65,"slug":66,"count":67,"latest_published_at":68},"Hardware","hardware",194,"2026-09-29T13:16:04.000Z",{"name":70,"slug":71,"count":72,"latest_published_at":18},"Science","science",155,{"name":74,"slug":75,"count":76,"latest_published_at":77},"Consumer Tech","consumer-tech",142,"2026-09-29T18:38:03.000Z",{"name":79,"slug":80,"count":81,"latest_published_at":82},"Software","software",91,"2026-09-25T20:55:00.000Z",{"name":84,"slug":85,"count":86,"latest_published_at":18},"Dev Tools","dev-tools",90,{"name":88,"slug":89,"count":90,"latest_published_at":91},"Startups","startups",83,"2026-09-29T21:51:36.000Z",{"name":93,"slug":94,"count":95,"latest_published_at":96},"General","general",49,"2026-09-28T16:44:57.000Z",{"name":98,"slug":99,"count":100,"latest_published_at":101},"Gaming","gaming",48,"2026-09-25T18:35:21.000Z",{"name":103,"slug":104,"count":105,"latest_published_at":106},"Reviews","reviews",31,"2026-09-28T14:31:34.000Z",{"name":108,"slug":109,"count":110,"latest_published_at":111},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]