[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-new-filter-catches-safety-blind-spots-in-multilingual-ai-models":10,"sections":35},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":24,"tags":25,"sources":30,"feedback":34,"feedback_at":22,"cost_usd":34,"total_tokens":34},7439,"new-filter-catches-safety-blind-spots-in-multilingual-ai-models","New Filter Catches Safety Blind Spots in Multilingual AI Models","New research finds multilingual AI safety signals span multiple layers, and a filter using all of them cuts harmful outputs by 60 percent.","A new filtering method shows that keeping large language models safe across languages takes more than checking one layer of the network.\n\nResearchers studying LLM fine-tuning have found that safety-degrading data can slip into training sets even when the data looks harmless, quietly eroding a model's safety alignment. Existing detection methods flag risky training examples by scanning activations in a single \"safety-sensitive\" layer, an approach built and tested mostly on English-only models. The team's cross-lingual analysis found that safety-relevant signals are not concentrated in one shared layer across languages. They are scattered across several layers, and only partially overlap between languages. Based on that finding, the researchers built MMSAFE, a multi-layer framework that scans several layers at once to catch both signals shared across languages and signals unique to just one.\n\nAcross multiple models, languages, and safety benchmarks, MMSAFE cut the rate of harmful responses by 60% compared with randomly filtering training data, and it beat the best single-layer detection method tested. That gap matters because most safety fine-tuning tooling has been built and validated in English, then assumed to generalize. This work shows that assumption does not hold, and safety filters designed around English-language models likely miss risks in other languages.\n\nCall it another case of \"safety-aligned\" quietly meaning \"safety-aligned in English.\" Closing that gap will take more than translating a benchmark. It means rethinking how the alignment signal itself behaves across languages.","[\"ai-safety\",\"multilingual-ai\",\"fine-tuning\",\"llm-research\"]","2026-09-23T04:00:00.000Z","2026-09-23T13:29:01.575Z","2026-09-23T13:29:07.313Z","published",null,[],"ai",[26,27,28,29],"ai-safety","multilingual-ai","fine-tuning","llm-research",[31],{"name":32,"url":33},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.22144",0,{"sections":36},[37,41,45,50,55,60,65,70,75,80,85,90,95,100],{"name":38,"slug":24,"count":39,"latest_published_at":40},"AI",4347,"2026-09-23T12:00:00.000Z",{"name":42,"slug":43,"count":44,"latest_published_at":18},"Security","security",713,{"name":46,"slug":47,"count":48,"latest_published_at":49},"Policy","policy",371,"2026-09-23T12:00:43.000Z",{"name":51,"slug":52,"count":53,"latest_published_at":54},"Deals","deals",212,"2026-09-23T13:00:46.000Z",{"name":56,"slug":57,"count":58,"latest_published_at":59},"Hardware","hardware",170,"2026-09-23T11:59:23.000Z",{"name":61,"slug":62,"count":63,"latest_published_at":64},"Science","science",134,"2026-09-23T09:00:00.000Z",{"name":66,"slug":67,"count":68,"latest_published_at":69},"Consumer Tech","consumer-tech",110,"2026-09-22T20:00:00.000Z",{"name":71,"slug":72,"count":73,"latest_published_at":74},"Software","software",81,"2026-09-23T09:56:13.000Z",{"name":76,"slug":77,"count":78,"latest_published_at":79},"Dev Tools","dev-tools",79,"2026-09-22T22:21:13.000Z",{"name":81,"slug":82,"count":83,"latest_published_at":84},"Startups","startups",65,"2026-09-22T22:06:48.000Z",{"name":86,"slug":87,"count":88,"latest_published_at":89},"Gaming","gaming",45,"2026-09-22T15:35:06.000Z",{"name":91,"slug":92,"count":93,"latest_published_at":94},"General","general",43,"2026-09-21T23:48:56.000Z",{"name":96,"slug":97,"count":98,"latest_published_at":99},"Reviews","reviews",27,"2026-09-22T13:00:00.000Z",{"name":101,"slug":102,"count":103,"latest_published_at":104},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]