A new benchmark says the hard part of letting AI write factory-floor code isn't the easy jobs, it's everything else.
Researchers built PLCWorld, a benchmark that pairs Structured Text execution with simulated plant responses to test whether LLM-written PLC programs, the code that runs robotic arms, conveyor belts, and sensors, actually do what they're supposed to without causing a safety violation. It includes 100 synthetic tasks and 473 task-condition pairs split between motion control and material handling, graded by how many interdependent control steps each task requires. The team validated the setup with practitioner review, reference solutions, and 542 deliberately broken counterexample programs, all of which triggered the evaluator's safety rules as designed. Testing GPT-5.5 directly, the researchers measured 82.70% task success on easy cases, falling to 25.10% on hard ones; five other LLMs and four generate-and-verify workflow variants showed similar drop-offs, with success, safety, and generation cost all trading off differently from model to model.
That's the real finding: a model can ace simple, isolated tests and still fail badly once it has to track multiple dependent control steps, which is exactly the kind of program that runs real equipment. PLCWorld tracks Task Success and Safety Violation as separate scores, and that distinction matters because a program can look functionally fine while still doing something risky to a machine, a split most coding benchmarks don't bother to make.
Nobody is proposing LLMs run your assembly line unsupervised yet, and on this evidence, nobody should.