A new review of agentic AI research says most testing checks what an AI agent does, not whether it stays safe over time.
The study is a systematic mapping review covering 262 papers on agent evaluation, software assurance, cyber-physical systems, runtime monitoring, and regulatory guidance. The authors sort existing validation work into five categories: behavioral, safety, temporal, regulatory, and multi-agent concerns. Behavioral evaluation dominates, appearing in 105 of the 262 papers, while temporal validity (whether an agent's decisions hold up as conditions change) shows up in just 14. Regulatory approaches lean mostly on assurance cases and regulatory analysis from IEEE-indexed venues, and multi-agent research relies mostly on benchmarks.
That imbalance matters because agentic systems plan, use tools, remember context, and adapt across multi-step trajectories, so a single input-output test says little about whether the system still behaves well ten steps later or after conditions shift. The paper backs this up with case studies in medical care, industrial operations, and smart-mobility systems, settings where an agent that looks fine in isolated testing can still drift into unsafe behavior over a longer run.
In other words, testing an AI agent like a function call misses the point. Agents fail in the middle of a trajectory, not just at the first or last step.