[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-open-weight-llms-show-big-refusal-gaps-for-somali-prompts":10,"sections":41},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":36,"feedback":40,"feedback_at":22,"cost_usd":40,"total_tokens":40},5086,"open-weight-llms-show-big-refusal-gaps-for-somali-prompts","Open-Weight LLMs Show Big Refusal Gaps for Somali Prompts","A new benchmark finds four open-weight models refuse harmful Somali prompts far less often than identical English ones, often failing incoherently instead.","Four open-weight language models are far more likely to refuse a harmful request in English than the identical request in Somali.\n\nResearchers evaluated Llama-3.1-8B-Instruct, Gemma-2-9B-Instruct, Qwen-2.5-7B-Instruct, and Aya-23-8B against SomaliBench v0, a native-author-verified set of 100 harmful-intent prompts paired in English and Somali. Each model ran locally at temperature 0 under the same English \"helpful, harmless, honest\" system prompt. All four showed a refusal gap between the English and Somali versions ranging from 0.40 to 0.93, and the gaps held up under paired bootstrap and exact McNemar significance tests. A pinned Claude Sonnet snapshot classified each response as refused, complied, or unclear; it declined to classify 34 of 800 responses, which the paper's native Somali-speaking author labeled by hand. In a separate check, that same author reviewed 74 comparable rows and agreed with the classifier's judgments 100% of the time, a Cohen's kappa of 1.00.\n\nFor three of the four models, the Somali refusals didn't fail by producing fluent harmful answers. They failed by producing wrong-language, incoherent, or off-topic text, a different problem than compliance but one a simple refusal-rate count can hide. A safety system that looks intact in English can be effectively absent in a language its developers never tested, and nobody monitoring English-only evals would catch it.\n\nEvery one of these models ships with the same English-language safety promises attached. This study is a reminder that almost nobody checks whether those promises survive translation; the researchers aren't releasing the raw generations, since some Somali outputs may contain harmful content.","[\"ai-safety\",\"low-resource-languages\",\"llm-evaluation\",\"open-weight-models\"]","2026-08-17T04:00:00.000Z","2026-08-17T09:46:42.403Z","2026-08-17T09:46:54.212Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"Remove the unsourced '20 million Somali speakers' figure and the insinuating spin on why transcripts aren't public (the source states plainly it's due to potential harmful content, not evasiveness), and stop conflating the 34-of-800 manually-labeled 'declined' classifications with the separate 74-row spot-check that actually produced the 100% agreement figure.","resolved","ai",[32,33,34,35],"ai-safety","low-resource-languages","llm-evaluation","open-weight-models",[37],{"name":38,"url":39},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2605.25420",0,{"sections":42},[43,47,51,56,61,66,71,76,81,86,91,96,101,106],{"name":44,"slug":30,"count":45,"latest_published_at":46},"AI",3293,"2026-08-20T04:00:00.000Z",{"name":48,"slug":49,"count":50,"latest_published_at":46},"Security","security",435,{"name":52,"slug":53,"count":54,"latest_published_at":55},"Policy","policy",210,"2026-08-19T09:32:27.000Z",{"name":57,"slug":58,"count":59,"latest_published_at":60},"Deals","deals",179,"2026-06-29T20:02:07.000Z",{"name":62,"slug":63,"count":64,"latest_published_at":65},"Hardware","hardware",140,"2026-08-19T18:25:42.000Z",{"name":67,"slug":68,"count":69,"latest_published_at":70},"Consumer Tech","consumer-tech",95,"2026-08-18T16:05:00.000Z",{"name":72,"slug":73,"count":74,"latest_published_at":75},"Science","science",90,"2026-08-19T18:41:02.000Z",{"name":77,"slug":78,"count":79,"latest_published_at":80},"Software","software",73,"2026-08-18T07:51:50.000Z",{"name":82,"slug":83,"count":84,"latest_published_at":85},"Dev Tools","dev-tools",69,"2026-08-18T04:00:00.000Z",{"name":87,"slug":88,"count":89,"latest_published_at":90},"Startups","startups",47,"2026-08-19T19:13:46.000Z",{"name":92,"slug":93,"count":94,"latest_published_at":95},"Gaming","gaming",41,"2026-07-09T04:00:00.000Z",{"name":97,"slug":98,"count":99,"latest_published_at":100},"General","general",33,"2026-08-18T22:18:13.000Z",{"name":102,"slug":103,"count":104,"latest_published_at":105},"Reviews","reviews",20,"2026-06-24T12:00:01.000Z",{"name":107,"slug":108,"count":109,"latest_published_at":110},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]