Large language models can write working Datalog code about two-thirds of the time on their own, and only reach the low eighties even with help.
A new paper, "DatalogBench: Evaluating Large Language Models on Text-to-Datalog Synthesis" (arXiv:2609.37233), tests six LLMs and four prompting setups on 136 tasks that ask a model to turn a plain-English question into a working Datalog program. Datalog is the logic-based language used for things like program analysis, and writing it by hand is notoriously fiddly. The paper's authors graded each generated program by actually running it against held-out inputs and comparing results to a verified oracle, not by eyeballing the code. Under straightforward prompting, exact-match accuracy topped out at 68.4%, and most failures happened before the code even ran, because models invented helper predicates they never declared. Note that this is an unreviewed preprint, not a peer-reviewed publication.
The bigger signal is what closes the gap and what doesn't. Two coding agents, which can iterate and check their own output, pushed accuracy to 83.8% and wiped out nearly all those compile-time errors, according to the paper's findings. But the errors that remained clustered in recursive tasks, suggesting today's models can pattern-match syntax without actually reasoning through recursive logic.
That's a familiar split for anyone watching AI coding tools: agents fix the sloppy mistakes, but the deep reasoning gaps stick around.