New research pins down when local AI models are safe to trust for network automation, and when they are not.
Researchers built Touchstone, a local-first pipeline that routes network automation queries to seven off-the-shelf small language models ranging from 1 billion to 8 billion parameters. Instead of trusting every answer outright, Touchstone runs each candidate through a task-specific intrinsic check, a cheap deterministic test that flags responses violating a necessary correctness condition. Anything that fails the check gets escalated to a frontier LLM instead of being used as-is. On conflict detection tasks the system hit 98.6% accuracy while escalating only 16% of queries, and on intent translation it reached 93.8% accuracy while escalating 17%.
The pitch is straightforward: sending production configs, topologies, and logs to a third-party frontier model is a data-exposure risk most network teams would rather avoid. Touchstone shows you can keep most of that traffic local without gutting accuracy, as long as the task has a way to verify its own output. That condition, not the accuracy numbers, is the real finding here.
On TeleQnA, a knowledge-only benchmark with no built-in way to check answers, Touchstone could not match the frontier baseline. The limit on local AI isn't model size. It's whether you can catch it being wrong.