[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-llms-still-struggle-to-write-datalog-new-benchmark-finds":10,"sections":40},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":35,"feedback":39,"feedback_at":22,"cost_usd":39,"total_tokens":39},8624,"llms-still-struggle-to-write-datalog-new-benchmark-finds","LLMs Still Struggle to Write Datalog, New Benchmark Finds","A new preprint finds LLMs peak at 68.4% accuracy translating plain English into Datalog, rising to 83.8% only with coding agents.","Large language models can write working Datalog code about two-thirds of the time on their own, and only reach the low eighties even with help.\n\nA new paper, \"DatalogBench: Evaluating Large Language Models on Text-to-Datalog Synthesis\" (arXiv:2609.37233), tests six LLMs and four prompting setups on 136 tasks that ask a model to turn a plain-English question into a working Datalog program. Datalog is the logic-based language used for things like program analysis, and writing it by hand is notoriously fiddly. The paper's authors graded each generated program by actually running it against held-out inputs and comparing results to a verified oracle, not by eyeballing the code. Under straightforward prompting, exact-match accuracy topped out at 68.4%, and most failures happened before the code even ran, because models invented helper predicates they never declared. Note that this is an unreviewed preprint, not a peer-reviewed publication.\n\nThe bigger signal is what closes the gap and what doesn't. Two coding agents, which can iterate and check their own output, pushed accuracy to 83.8% and wiped out nearly all those compile-time errors, according to the paper's findings. But the errors that remained clustered in recursive tasks, suggesting today's models can pattern-match syntax without actually reasoning through recursive logic.\n\nThat's a familiar split for anyone watching AI coding tools: agents fix the sloppy mistakes, but the deep reasoning gaps stick around.","[\"ai\",\"benchmarks\",\"datalog\",\"research\"]","2026-09-30T04:00:00.000Z","2026-09-30T15:38:11.443Z","2026-09-30T15:38:17.964Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"Attribute the findings explicitly to the source paper (title 'DatalogBench: Evaluating Large Language Models on Text-to-Datalog Synthesis,' arXiv:2609.37233) and note it's an unreviewed preprint, rather than the vague 'researchers built DatalogBench.'","resolved","ai",[30,32,33,34],"benchmarks","datalog","research",[36],{"name":37,"url":38},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.37233",0,{"sections":41},[42,45,49,53,58,63,68,73,78,82,87,92,97,102],{"name":43,"slug":30,"count":44,"latest_published_at":18},"AI",5147,{"name":46,"slug":47,"count":48,"latest_published_at":18},"Security","security",788,{"name":50,"slug":51,"count":52,"latest_published_at":18},"Policy","policy",417,{"name":54,"slug":55,"count":56,"latest_published_at":57},"Deals","deals",284,"2026-09-29T21:00:00.000Z",{"name":59,"slug":60,"count":61,"latest_published_at":62},"Hardware","hardware",194,"2026-09-29T13:16:04.000Z",{"name":64,"slug":65,"count":66,"latest_published_at":67},"Science","science",154,"2026-09-28T13:19:18.000Z",{"name":69,"slug":70,"count":71,"latest_published_at":72},"Consumer Tech","consumer-tech",142,"2026-09-29T18:38:03.000Z",{"name":74,"slug":75,"count":76,"latest_published_at":77},"Software","software",91,"2026-09-25T20:55:00.000Z",{"name":79,"slug":80,"count":81,"latest_published_at":18},"Dev Tools","dev-tools",90,{"name":83,"slug":84,"count":85,"latest_published_at":86},"Startups","startups",83,"2026-09-29T21:51:36.000Z",{"name":88,"slug":89,"count":90,"latest_published_at":91},"General","general",49,"2026-09-28T16:44:57.000Z",{"name":93,"slug":94,"count":95,"latest_published_at":96},"Gaming","gaming",48,"2026-09-25T18:35:21.000Z",{"name":98,"slug":99,"count":100,"latest_published_at":101},"Reviews","reviews",31,"2026-09-28T14:31:34.000Z",{"name":103,"slug":104,"count":105,"latest_published_at":106},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]