[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-skynet-tool-flags-ai-agent-workflows-before-failures-spread":10,"sections":44},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":34,"tags":35,"sources":39,"feedback":43,"feedback_at":22,"cost_usd":43,"total_tokens":43},6227,"skynet-tool-flags-ai-agent-workflows-before-failures-spread","Skynet Tool Flags AI Agent Workflows Before Failures Spread","Skynet turns AI agent workflows into graphs to catch failures and attacks that single-step checks miss.","A new tool inspects how AI agents work together as a whole, not just step by step, to catch failures before they spread.\n\nResearchers describe Skynet, a workflow-level anomaly detection framework for agentic AI, in a paper posted to arXiv this week. Rather than inspecting individual prompts or tool calls in isolation, Skynet converts an AI system's full execution, its planning steps, tool invocations, and multi-agent coordination, into a directed graph, then checks that graph's structure and semantic content against a model built purely from benign runs. The system trains only on workflows that behaved correctly. Any execution whose graph looks structurally or semantically off, whether triggered by a flawed plan or an injected prompt, gets flagged, even for failure types the training data never included.\n\nThat distinction matters because agentic systems tend to fail in ways local checks miss: one corrupted step early in a workflow propagates through every downstream agent and tool call it touches, and the resulting inconsistency only becomes visible when you look at the whole dependency chain. Training on benign behavior alone also means Skynet does not need labeled examples of every possible attack, a real advantage as the number of ways agentic systems can be hijacked keeps growing.\n\nOn three public agentic safety and failure benchmarks, the system reportedly held false positives under 1% while catching most real failures, with latency low enough for live monitoring. Three benchmarks are a narrow slice of how agentic systems actually get deployed, though, and their tool ecosystems and coordination patterns vary widely. Wider adoption will depend on whether the graph model holds up against workflows and integrations the benchmarks never covered.","[\"ai\",\"agentic-ai\",\"security\",\"research\"]","2026-09-10T04:00:00.000Z","2026-09-10T08:04:26.254Z","2026-09-10T08:04:38.144Z","published",null,[24,30],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"The headline and dek assert the researchers are from Google, but nothing in the source material establishes any Google affiliation — remove or verify the attribution before republishing.","resolved",{"id":31,"reviewer":26,"round":32,"reason":33,"status":29},"editor-r2",2,"The Google attribution issue is fixed, but the closing paragraph is speculative editorializing about the benchmark's name and doesn't end on a concrete concluding statement — replace it with a substantive close (e.g., limitations of the three benchmarks, or what would need to happen for wider adoption).","ai",[34,36,37,38],"agentic-ai","security","research",[40],{"name":41,"url":42},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.06835",0,{"sections":45},[46,50,53,58,63,68,73,77,82,87,92,97,102,107],{"name":47,"slug":34,"count":48,"latest_published_at":49},"AI",3480,"2026-09-11T04:00:00.000Z",{"name":51,"slug":37,"count":52,"latest_published_at":49},"Security",628,{"name":54,"slug":55,"count":56,"latest_published_at":57},"Policy","policy",336,"2026-09-11T00:56:21.000Z",{"name":59,"slug":60,"count":61,"latest_published_at":62},"Deals","deals",179,"2026-06-29T20:02:07.000Z",{"name":64,"slug":65,"count":66,"latest_published_at":67},"Hardware","hardware",153,"2026-09-09T15:12:32.000Z",{"name":69,"slug":70,"count":71,"latest_published_at":72},"Consumer Tech","consumer-tech",99,"2026-09-09T17:27:33.000Z",{"name":74,"slug":75,"count":76,"latest_published_at":49},"Science","science",98,{"name":78,"slug":79,"count":80,"latest_published_at":81},"Software","software",75,"2026-09-10T20:41:21.000Z",{"name":83,"slug":84,"count":85,"latest_published_at":86},"Dev Tools","dev-tools",69,"2026-08-18T04:00:00.000Z",{"name":88,"slug":89,"count":90,"latest_published_at":91},"Startups","startups",55,"2026-09-09T23:14:29.000Z",{"name":93,"slug":94,"count":95,"latest_published_at":96},"Gaming","gaming",43,"2026-09-10T12:18:06.000Z",{"name":98,"slug":99,"count":100,"latest_published_at":101},"General","general",41,"2026-09-08T01:57:23.000Z",{"name":103,"slug":104,"count":105,"latest_published_at":106},"Reviews","reviews",20,"2026-06-24T12:00:01.000Z",{"name":108,"slug":109,"count":110,"latest_published_at":111},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]