[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-watermarking-text-can-make-ai-models-ignore-safety-rules":10,"sections":41},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":36,"feedback":40,"feedback_at":22,"cost_usd":40,"total_tokens":40},6990,"watermarking-text-can-make-ai-models-ignore-safety-rules","Watermarking Text Can Make AI Models Ignore Safety Rules","New research finds that watermarking AI text with SynthID-Text can change word choices and weaken safety guardrails under adversarial prompts.","Anthropic's plan to watermark Claude's output has a side effect: it can make models more likely to break their own safety rules.\n\nAnthropic recently disclosed that future Claude models will use SynthID-Text, an open-source watermarking method built by Google, to comply with a new EU disclosure law. The system uses a secret key to nudge word choices, swapping \"cloudy\" for \"overcast,\" for example, so anyone holding the key can later verify a piece of text came from that model. New research shows the technique does more than tweak vocabulary. It can also change which tools a model calls and how consistently it follows its safety training, with the effect growing sharper under adversarial prompts designed to trick a model into doing something it normally wouldn't, like revealing a password.\n\nThat's a bigger deal than it sounds. Watermarking is usually pitched as invisible bookkeeping, a way to prove text is AI-generated without changing what the text says. Lasso Security researcher Andrea Siposova said that altering how a model generates text \"is definitely going to change their behavior, especially when we place it under adversarial conditions,\" and that such tradeoffs \"show up somewhere\" even when the watermark itself isn't perceptible to a reader.\n\nThe irony: a compliance feature built to satisfy one regulation may be quietly undermining the safety testing meant to satisfy another.","[\"ai safety\",\"watermarking\",\"anthropic\",\"security research\"]","2026-09-17T18:33:13.000Z","2026-09-19T02:06:06.246Z","2026-09-19T02:06:18.160Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"Cut or attribute the unsupported speculation that 'other regions expected to follow' and that 'every major lab' now has to bolt on watermarking — neither claim is in the source material and both go beyond what Lasso's research or Anthropic's disclosure actually support.","resolved","security",[32,33,34,35],"ai safety","watermarking","anthropic","security research",[37],{"name":38,"url":39},"Ars Technica","https:\u002F\u002Farstechnica.com\u002Fsecurity\u002F2026\u002F09\u002Fai-text-watermarking-can-make-models-more-vulnerable-to-adversarial-prompts\u002F",0,{"sections":42},[43,48,52,57,62,66,71,76,80,85,90,95,100,105],{"name":44,"slug":45,"count":46,"latest_published_at":47},"AI","ai",4083,"2026-09-18T04:00:00.000Z",{"name":49,"slug":30,"count":50,"latest_published_at":51},"Security",663,"2026-09-18T14:00:14.000Z",{"name":53,"slug":54,"count":55,"latest_published_at":56},"Policy","policy",339,"2026-09-17T12:00:00.000Z",{"name":58,"slug":59,"count":60,"latest_published_at":61},"Deals","deals",179,"2026-06-29T20:02:07.000Z",{"name":63,"slug":64,"count":65,"latest_published_at":47},"Hardware","hardware",155,{"name":67,"slug":68,"count":69,"latest_published_at":70},"Science","science",126,"2026-09-18T14:41:02.000Z",{"name":72,"slug":73,"count":74,"latest_published_at":75},"Consumer Tech","consumer-tech",99,"2026-09-09T17:27:33.000Z",{"name":77,"slug":78,"count":79,"latest_published_at":47},"Dev Tools","dev-tools",78,{"name":81,"slug":82,"count":83,"latest_published_at":84},"Software","software",75,"2026-09-10T20:41:21.000Z",{"name":86,"slug":87,"count":88,"latest_published_at":89},"Startups","startups",55,"2026-09-09T23:14:29.000Z",{"name":91,"slug":92,"count":93,"latest_published_at":94},"Gaming","gaming",43,"2026-09-10T12:18:06.000Z",{"name":96,"slug":97,"count":98,"latest_published_at":99},"General","general",41,"2026-09-08T01:57:23.000Z",{"name":101,"slug":102,"count":103,"latest_published_at":104},"Reviews","reviews",20,"2026-06-24T12:00:01.000Z",{"name":106,"slug":107,"count":108,"latest_published_at":109},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]