[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-benchmark-finds-malicious-ai-skill-detectors-struggle-on-new-sources":10,"sections":41},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":36,"feedback":40,"feedback_at":22,"cost_usd":40,"total_tokens":40},5774,"benchmark-finds-malicious-ai-skill-detectors-struggle-on-new-sources","Benchmark Finds Malicious AI Skill Detectors Struggle on New Sources","A new benchmark of over 9,700 AI Skill packages shows malicious-code detectors' Macro-F1 scores drop from 93.2% to 66.5% when tested on unseen sources.","A new benchmark for catching malicious AI Agent Skills shows today's detectors mostly work only on threats they have already seen.\n\nMaliciousSkillBench pools 13 public datasets of Agent Skills, the reusable instruction packages, scripts, and configs that extend LLM agents, then dedupes 8,414 raw malicious entries into 7,539 unique identities and lands on a final set of 9,740 Skills: 7,505 malicious and 2,235 benign across 11 harmonized attack categories. The researchers tested three learned text classifiers and three off-the-shelf Skill scanners. The best classifier, a word-level TF-IDF model paired with an SVM, hit a 93.2% Macro-F1 score on a standard random split, but that fell to 66.5% when evaluated on sources excluded from training, all while still flagging 62.4 percent of benign Skills as malicious. Off-the-shelf scanners fared no better, trading recall for lower false-positive rates instead of improving on both.\n\nThat gap matters because real attackers rarely publish through the same channels a detector was trained on. A tool that looks reliable on familiar data can still miss most new threats, or drown genuine alerts in false positives, the same failure mode that has dogged npm and browser-extension malware scanners for years. As Skill marketplaces expand, that source-disjoint blind spot is exactly where attackers will aim.\n\nA scanner that only catches malware it has already met isn't a detector, it's a memory test.","[\"ai-security\",\"malicious-skills\",\"malware-detection\",\"llm-agents\"]","2026-08-21T04:00:00.000Z","2026-08-21T05:29:56.060Z","2026-08-21T05:30:07.158Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"Fix two mismatches with the source: the piece labels 93.2%\u002F66.5% as 'accuracy' when the paper explicitly reports these as Macro-F1 scores, and the dek's claim that detectors 'lose most of their accuracy' overstates a source that shows only a ~29% relative decline (93.2 to 66.5), not a majority loss.","resolved","security",[32,33,34,35],"ai-security","malicious-skills","malware-detection","llm-agents",[37],{"name":38,"url":39},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2608.19901",0,{"sections":42},[43,47,50,55,60,65,70,75,80,85,90,95,100,105],{"name":44,"slug":45,"count":46,"latest_published_at":18},"AI","ai",3300,{"name":48,"slug":30,"count":49,"latest_published_at":18},"Security",449,{"name":51,"slug":52,"count":53,"latest_published_at":54},"Policy","policy",211,"2026-08-20T10:47:43.000Z",{"name":56,"slug":57,"count":58,"latest_published_at":59},"Deals","deals",179,"2026-06-29T20:02:07.000Z",{"name":61,"slug":62,"count":63,"latest_published_at":64},"Hardware","hardware",141,"2026-08-20T11:20:00.000Z",{"name":66,"slug":67,"count":68,"latest_published_at":69},"Consumer Tech","consumer-tech",95,"2026-08-18T16:05:00.000Z",{"name":71,"slug":72,"count":73,"latest_published_at":74},"Science","science",91,"2026-08-20T10:01:48.000Z",{"name":76,"slug":77,"count":78,"latest_published_at":79},"Software","software",73,"2026-08-18T07:51:50.000Z",{"name":81,"slug":82,"count":83,"latest_published_at":84},"Dev Tools","dev-tools",69,"2026-08-18T04:00:00.000Z",{"name":86,"slug":87,"count":88,"latest_published_at":89},"Startups","startups",48,"2026-08-20T18:34:26.000Z",{"name":91,"slug":92,"count":93,"latest_published_at":94},"Gaming","gaming",41,"2026-07-09T04:00:00.000Z",{"name":96,"slug":97,"count":98,"latest_published_at":99},"General","general",33,"2026-08-18T22:18:13.000Z",{"name":101,"slug":102,"count":103,"latest_published_at":104},"Reviews","reviews",20,"2026-06-24T12:00:01.000Z",{"name":106,"slug":107,"count":108,"latest_published_at":109},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]