[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-finetune-an-ai-on-narrow-ethics-and-it-generalizes":10,"sections":34},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":24,"tags":25,"sources":29,"feedback":33,"feedback_at":22,"cost_usd":33,"total_tokens":33},9233,"finetune-an-ai-on-narrow-ethics-and-it-generalizes","Finetune an AI on narrow ethics, and it generalizes","New research shows that finetuning an AI model on narrow ethical rules can make it safer broadly, but how well that generalizes varies a lot by approach.","Train an AI model on a narrow slice of ethics, and it can start behaving more aligned everywhere else.\n\nA new paper finetunes a helpfulness-only language model using the Constitutional AI method, built on four distinct ethical frameworks: deontology, consequentialism, virtue ethics, and a framework that treats the AI as subordinate to human authority. For each framework, researchers trained on just two narrow safety subcategories, then tested the resulting model against a broad set of general safety categories, including ones deliberately excluded from the training data. The narrow finetuning reliably produced what the authors call \"emergent alignment\": safer behavior well beyond what the model was directly taught. A separate, more granular \"ethical persona\" test confirmed the models absorbed the worldview baked into their constitution, with the consequentialist-trained model, for instance, agreeing more with utilitarian statements than deontological ones.\n\nThis flips a darker finding from earlier \"emergent misalignment\" research, where finetuning on narrow bad behavior produced broadly bad behavior. Together, the two results support the idea that alignment training works less like rule memorization and more like selecting a character for the model to play, a theory the paper calls the \"persona selection\" hypothesis. The catch: how well that persona held up varied a lot depending on which constitution and which finetuning approach was used.\n\nThat inconsistency is the real headline. A model can pass its ethics exam in the lab and still answer unpredictably once the question changes.","[\"ai-alignment\",\"llm-safety\",\"constitutional-ai\"]","2026-10-01T04:00:00.000Z","2026-10-02T03:13:24.849Z","2026-10-02T03:13:28.779Z","published",null,[],"ai",[26,27,28],"ai-alignment","llm-safety","constitutional-ai",[30],{"name":31,"url":32},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2606.09475",0,{"sections":35},[36,39,43,47,52,57,61,66,71,75,80,85,90,95],{"name":37,"slug":24,"count":38,"latest_published_at":18},"AI",5629,{"name":40,"slug":41,"count":42,"latest_published_at":18},"Security","security",816,{"name":44,"slug":45,"count":46,"latest_published_at":18},"Policy","policy",430,{"name":48,"slug":49,"count":50,"latest_published_at":51},"Deals","deals",298,"2026-09-30T21:00:26.000Z",{"name":53,"slug":54,"count":55,"latest_published_at":56},"Hardware","hardware",196,"2026-09-30T13:00:00.000Z",{"name":58,"slug":59,"count":60,"latest_published_at":18},"Science","science",163,{"name":62,"slug":63,"count":64,"latest_published_at":65},"Consumer Tech","consumer-tech",149,"2026-09-30T22:57:11.000Z",{"name":67,"slug":68,"count":69,"latest_published_at":70},"Dev Tools","dev-tools",93,"2026-10-01T02:30:48.000Z",{"name":72,"slug":73,"count":69,"latest_published_at":74},"Software","software","2026-09-30T21:41:11.000Z",{"name":76,"slug":77,"count":78,"latest_published_at":79},"Startups","startups",84,"2026-09-30T20:39:09.000Z",{"name":81,"slug":82,"count":83,"latest_published_at":84},"Gaming","gaming",51,"2026-09-30T16:24:30.000Z",{"name":86,"slug":87,"count":88,"latest_published_at":89},"General","general",50,"2026-09-30T21:37:54.000Z",{"name":91,"slug":92,"count":93,"latest_published_at":94},"Reviews","reviews",31,"2026-09-28T14:31:34.000Z",{"name":96,"slug":97,"count":98,"latest_published_at":99},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]