[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-ai-models-detect-harmful-memes-internally-but-fail-to-act":10,"sections":35},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":24,"tags":25,"sources":30,"feedback":34,"feedback_at":22,"cost_usd":34,"total_tokens":34},6640,"ai-models-detect-harmful-memes-internally-but-fail-to-act","AI Models Detect Harmful Memes Internally but Fail to Act","New research on Gemma-3 and Qwen3.5 shows the models often have the right answer buried inside them, but their outputs can't reliably tap it.","Vision-language models often 'know' a meme is hateful somewhere inside their weights - they just can't get that knowledge to their final answer.\n\nResearchers tested Gemma-3 and Qwen3.5 using sparse autoencoders, role-conditioned probes, causal interventions, and recovery experiments across six harmful-content benchmarks, including Spanish and Hindi-English code-mixed cases. They found that reading the internal sparse features produced far better harmful-meme classification than the models' own native outputs: Qwen's sparse readout averaged 0.740 macro-F1 versus just 0.432 natively, and Gemma jumped from 0.532 to 0.714. A case study on Gemma-3-12B with Facebook's Hateful Memes dataset found a rank-32 image-prompt interaction pattern that scored 0.756 versus the model's native 0.685 macro-F1.\n\nThis points to a different bottleneck than the usual explanation for moderation failures. The models aren't necessarily missing the evidence that a meme is harmful - they fail to route that evidence to the output layer. The team showed calibration-only routing fixes recovered 93.3 percent of that accuracy gap without retraining the whole model.\n\nStill, the fix isn't free. Distilling the probes into LoRA adapters helped on individual tasks but caused negative transfer when shared across tasks - patching one blind spot opened another.","[\"ai interpretability\",\"harmful content detection\",\"vision-language models\",\"ai safety\"]","2026-09-17T04:00:00.000Z","2026-09-18T03:55:56.122Z","2026-09-18T03:56:08.066Z","published",null,[],"ai",[26,27,28,29],"ai interpretability","harmful content detection","vision-language models","ai safety",[31],{"name":32,"url":33},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.18860",0,{"sections":36},[37,41,45,50,55,59,63,68,73,77,82,87,92,97],{"name":38,"slug":24,"count":39,"latest_published_at":40},"AI",3852,"2026-09-17T08:27:09.000Z",{"name":42,"slug":43,"count":44,"latest_published_at":18},"Security","security",648,{"name":46,"slug":47,"count":48,"latest_published_at":49},"Policy","policy",338,"2026-09-11T04:00:00.000Z",{"name":51,"slug":52,"count":53,"latest_published_at":54},"Deals","deals",179,"2026-06-29T20:02:07.000Z",{"name":56,"slug":57,"count":58,"latest_published_at":18},"Hardware","hardware",154,{"name":60,"slug":61,"count":62,"latest_published_at":18},"Science","science",114,{"name":64,"slug":65,"count":66,"latest_published_at":67},"Consumer Tech","consumer-tech",99,"2026-09-09T17:27:33.000Z",{"name":69,"slug":70,"count":71,"latest_published_at":72},"Software","software",75,"2026-09-10T20:41:21.000Z",{"name":74,"slug":75,"count":76,"latest_published_at":18},"Dev Tools","dev-tools",73,{"name":78,"slug":79,"count":80,"latest_published_at":81},"Startups","startups",55,"2026-09-09T23:14:29.000Z",{"name":83,"slug":84,"count":85,"latest_published_at":86},"Gaming","gaming",43,"2026-09-10T12:18:06.000Z",{"name":88,"slug":89,"count":90,"latest_published_at":91},"General","general",41,"2026-09-08T01:57:23.000Z",{"name":93,"slug":94,"count":95,"latest_published_at":96},"Reviews","reviews",20,"2026-06-24T12:00:01.000Z",{"name":98,"slug":99,"count":100,"latest_published_at":101},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]