[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-new-benchmark-finds-arabic-ai-fails-half-of-safety-tests":10,"sections":41},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":36,"feedback":40,"feedback_at":22,"cost_usd":40,"total_tokens":40},5840,"new-benchmark-finds-arabic-ai-fails-half-of-safety-tests","New Benchmark Finds Arabic AI Fails Half of Safety Tests","A new human-curated Arabic redteaming benchmark shows leading language models, including GPT-4o and Claude 3.7 Sonnet, miss half of unsafe prompts.","Most leading AI models can't reliably refuse unsafe requests in Arabic, according to a new benchmark built specifically to test them.\n\nResearchers introduced ASAS, a human-curated benchmark of 801 adversarial prompts covering 8 safety categories and 8 attack strategies, designed to redteam large language models in Modern Standard Arabic. They tested seven models with Arabic capabilities, including GPT-4o, Claude 3.7 Sonnet, and regional models ALLaM and FANAR, then had human annotators score every response on a 4-point safety scale. Most models failed to block at least half of the unsafe prompts thrown at them. The weakest spots were high-harm topics like weapons and illicit substances, and the most effective attacks were the simplest: asking directly or lightly disguising the request.\n\nThis suggests the safety training these models get in English doesn't carry over cleanly to Arabic, even for models marketed as multilingual. It also calls into question a common shortcut in AI safety testing: using another model, like GPT-4o, as an automated judge. The researchers found those automated judges performed worse than human reviewers at catching unsafe responses.\n\nFor a region where these models are increasingly deployed in customer service, education, and government tools, that gap between English-language safety marketing and actual Arabic-language behavior is the kind of detail worth checking before, not after, deployment.","[\"ai-safety\",\"arabic-nlp\",\"llm-benchmarks\",\"redteaming\"]","2026-08-25T04:00:00.000Z","2026-08-25T04:53:52.516Z","2026-08-25T04:54:04.410Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"publisher-r1","publisher",1,"The benchmark's acronym 'ASAS' doesn't match its stated full name 'the Arabic Safety Index' (which would abbreviate to ASI), an internal naming inconsistency that needs correction before publishing.","resolved","ai",[32,33,34,35],"ai-safety","arabic-nlp","llm-benchmarks","redteaming",[37],{"name":38,"url":39},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2608.21985",0,{"sections":42},[43,46,50,55,60,65,70,75,80,85,90,95,100,105],{"name":44,"slug":30,"count":45,"latest_published_at":18},"AI",3329,{"name":47,"slug":48,"count":49,"latest_published_at":18},"Security","security",468,{"name":51,"slug":52,"count":53,"latest_published_at":54},"Policy","policy",224,"2026-08-24T19:02:32.000Z",{"name":56,"slug":57,"count":58,"latest_published_at":59},"Deals","deals",179,"2026-06-29T20:02:07.000Z",{"name":61,"slug":62,"count":63,"latest_published_at":64},"Hardware","hardware",145,"2026-08-22T21:25:33.000Z",{"name":66,"slug":67,"count":68,"latest_published_at":69},"Consumer Tech","consumer-tech",95,"2026-08-18T16:05:00.000Z",{"name":71,"slug":72,"count":73,"latest_published_at":74},"Science","science",91,"2026-08-20T10:01:48.000Z",{"name":76,"slug":77,"count":78,"latest_published_at":79},"Software","software",73,"2026-08-18T07:51:50.000Z",{"name":81,"slug":82,"count":83,"latest_published_at":84},"Dev Tools","dev-tools",69,"2026-08-18T04:00:00.000Z",{"name":86,"slug":87,"count":88,"latest_published_at":89},"Startups","startups",51,"2026-08-24T13:47:26.000Z",{"name":91,"slug":92,"count":93,"latest_published_at":94},"Gaming","gaming",41,"2026-07-09T04:00:00.000Z",{"name":96,"slug":97,"count":98,"latest_published_at":99},"General","general",33,"2026-08-18T22:18:13.000Z",{"name":101,"slug":102,"count":103,"latest_published_at":104},"Reviews","reviews",20,"2026-06-24T12:00:01.000Z",{"name":106,"slug":107,"count":108,"latest_published_at":109},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]