A new study says small, locally run AI agents can handle real data engineering work, if you let them see their own mistakes.
Researchers built a benchmark of fifteen mobility-data-workflow tasks covering data discovery, connectors, transit-feed processing, semantic enrichment, feature engineering, validation, visualization, and reporting. They tested ten local model configurations in both one-shot and closed-loop modes, five repetitions each, for 1,500 scored attempts on a single consumer-grade GPU. Deterministic checkers graded the actual output artifacts - scripts, tables, files, figures, reports - rather than just whether the text sounded plausible. The closed-loop setup, where an agent can inspect and repair its own intermediate output, lifted pass rates by 26.7 to 52.0 percentage points over single-shot generation for models above two billion parameters, with the best configuration hitting 85.3% success and a quantized 9-billion-parameter model reaching 69.3% on roughly 6.5 GB of memory.
The real finding here isn't "AI agent does good job" - it's that verifiability, not raw model size, is what makes local agents usable for this kind of work. That's a direct argument for running smaller open-weight models on your own hardware instead of routing every data task to a hosted frontier model, as long as the workflow exposes its own errors along the way.
Mobility pipelines are a tidy, well-structured test case; whether this closed-loop trick holds up on messier, less deterministic engineering work is the question the paper leaves open.