AI/ ai agents · open-weight models · data engineering · benchmarks

Local AI Agents Can Do Real Data Engineering, Study Finds

A new benchmark shows open-weight AI agents can finish most data engineering tasks locally, but only with a feedback loop to catch their own errors.

A new study says small, locally run AI agents can handle real data engineering work, if you let them see their own mistakes.

Researchers built a benchmark of fifteen mobility-data-workflow tasks covering data discovery, connectors, transit-feed processing, semantic enrichment, feature engineering, validation, visualization, and reporting. They tested ten local model configurations in both one-shot and closed-loop modes, five repetitions each, for 1,500 scored attempts on a single consumer-grade GPU. Deterministic checkers graded the actual output artifacts - scripts, tables, files, figures, reports - rather than just whether the text sounded plausible. The closed-loop setup, where an agent can inspect and repair its own intermediate output, lifted pass rates by 26.7 to 52.0 percentage points over single-shot generation for models above two billion parameters, with the best configuration hitting 85.3% success and a quantized 9-billion-parameter model reaching 69.3% on roughly 6.5 GB of memory.

The real finding here isn't "AI agent does good job" - it's that verifiability, not raw model size, is what makes local agents usable for this kind of work. That's a direct argument for running smaller open-weight models on your own hardware instead of routing every data task to a hosted frontier model, as long as the workflow exposes its own errors along the way.

Mobility pipelines are a tidy, well-structured test case; whether this closed-loop trick holds up on messier, less deterministic engineering work is the question the paper leaves open.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →