A new study turns 150 real AI system meltdowns into a field guide for not repeating them.
Researchers analyzed 150 production incident reports from open-source compound AI projects and anonymized enterprise deployments, building a taxonomy of 23 failure modes across five categories: retrieval, generation, tool, orchestration, and integration failures. They tested fixes with controlled fault injection experiments. Circuit breakers alone cut cascade propagation - errors spreading from one component to another - by 89%. Output quality gates caught 73% of silent quality degradation before it reached users, and isolating components cut a failure's blast radius by 64%.
The 71% drop in mean-time-to-recovery only showed up in systems stacking three or more of these patterns together, not from any single fix. That is the real finding: resilience in multi-component AI systems is additive, not a silver bullet. Compound AI systems - chains of retrieval steps, tools, and models - fail at the seams between components, not inside any one model, which is exactly where most teams aren't looking.
Call it the DevOps lesson tech keeps relearning: monitoring a system's parts tells you nothing about how they fail together.