A new assurance-focused review says the military's testing playbook for agentic AI can't actually prove those systems will behave safely once deployed.
Researchers examined 240 documented testing and evaluation practices across eight evaluation dimensions and three lifecycle stages for agentic AI used in military command and control. They found eight assumptions baked into standard testing methods, grouped into four clusters - specifiability, stability, composability, and supervisability - that assume a system's behavior is fixed and predictable enough to test in the first place. Agentic AI, which acts with more autonomy and adapts its behavior in ways traditional software doesn't, weakens every one of those assumptions. The paper argues this doesn't invalidate the test results themselves, but it breaks the logical link between "this system passed testing" and "this system will act the same way in the field."
The military has made public commitments to rigorous testing and human oversight before fielding AI in command and control, where the stakes are literally life and death. This paper says that even a system that clears every procedural box hasn't necessarily earned trust for real operations, because passing today's tests says less than usual about tomorrow's behavior when the system itself can adapt. Some narrower claims - about bounded mission limits, runtime constraints, and how much a system's output varies run to run - are still achievable in principle, but only as testing methods mature.
The paper's proposed fix, pushing part of the safety burden into deployment itself with expiration dates on unproven assumptions and a named owner for each one, amounts to admitting that "tested and approved" for agentic AI may need to become an ongoing job rather than a one-time stamp.