[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-budget-ai-models-get-java-coding-tasks-wrong-most-of-the-time":10,"sections":40},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":35,"feedback":39,"feedback_at":22,"cost_usd":39,"total_tokens":39},6594,"budget-ai-models-get-java-coding-tasks-wrong-most-of-the-time","Budget AI Models Get Java Coding Tasks Wrong Most of the Time","An arXiv preprint found three cheap LLMs correctly solved Java coding tasks just 12.9% of the time, often returning answers they never computed.","A new arXiv preprint says three budget-tier language models write Java code that compiles fine and barely works.\n\nThe study (arXiv:2609.18052) had Gemini Flash 3, GPT-5.4 mini and Claude Haiku 4.5 each write 992 algorithmic problems as Java Spring Boot service methods, following a fixed method signature and data-transfer-object spec. The preprint's authors ran four model-and-coding-tool pairings against two prompt styles, for eight configurations total, and forbade iteration or hardcoded answers. The resulting 7,593 methods were sorted into eight outcome categories, from methods that computed the right answer to ones that returned something without computing it, then deployed and actually run, producing 7,936 measured requests. The models nailed the required method signatures almost every time; what they put inside those methods was another matter.\n\n38.4% of submitted methods returned a value without computing it: stubs, defaults, or code that looked plausible but skipped the actual work. Only 12.9% of all answers were correct, and even the methods that did attempt real computation were right just 19.3% of the time. That gap between code that looks finished and code that actually works matters a lot if you are piping model output straight into a production service.\n\nThe paper's own authors call the results provisional, citing single-run testing, partial harness coverage and likely training-data contamination. Still, the pattern is a familiar one: the cheap model is cheap for a reason.","[\"ai\",\"llm-benchmarks\",\"coding\",\"java\"]","2026-09-17T04:00:00.000Z","2026-09-18T01:44:58.491Z","2026-09-18T01:45:10.426Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"Attribute the study explicitly (e.g., cite the arXiv preprint\u002Fits identifier) instead of the unnamed 'Researchers tested...' framing, since the piece never names the publication or institution behind the benchmark.","resolved","ai",[30,32,33,34],"llm-benchmarks","coding","java",[36],{"name":37,"url":38},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.18052",0,{"sections":41},[42,46,50,55,60,64,68,73,78,82,87,92,97,102],{"name":43,"slug":30,"count":44,"latest_published_at":45},"AI",3852,"2026-09-17T08:27:09.000Z",{"name":47,"slug":48,"count":49,"latest_published_at":18},"Security","security",648,{"name":51,"slug":52,"count":53,"latest_published_at":54},"Policy","policy",338,"2026-09-11T04:00:00.000Z",{"name":56,"slug":57,"count":58,"latest_published_at":59},"Deals","deals",179,"2026-06-29T20:02:07.000Z",{"name":61,"slug":62,"count":63,"latest_published_at":18},"Hardware","hardware",154,{"name":65,"slug":66,"count":67,"latest_published_at":18},"Science","science",114,{"name":69,"slug":70,"count":71,"latest_published_at":72},"Consumer Tech","consumer-tech",99,"2026-09-09T17:27:33.000Z",{"name":74,"slug":75,"count":76,"latest_published_at":77},"Software","software",75,"2026-09-10T20:41:21.000Z",{"name":79,"slug":80,"count":81,"latest_published_at":18},"Dev Tools","dev-tools",73,{"name":83,"slug":84,"count":85,"latest_published_at":86},"Startups","startups",55,"2026-09-09T23:14:29.000Z",{"name":88,"slug":89,"count":90,"latest_published_at":91},"Gaming","gaming",43,"2026-09-10T12:18:06.000Z",{"name":93,"slug":94,"count":95,"latest_published_at":96},"General","general",41,"2026-09-08T01:57:23.000Z",{"name":98,"slug":99,"count":100,"latest_published_at":101},"Reviews","reviews",20,"2026-06-24T12:00:01.000Z",{"name":103,"slug":104,"count":105,"latest_published_at":106},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]