AI/ ai-safety · red-teaming · reinforcement-learning · language-models

AI Red Teamers Learn to Attack and Defend Using GRPO

A new RL method pits attacker and defender models against each other, using GDPO to blend judge rewards so no single signal dominates.

Researchers taught one AI to jailbreak another, then made both sides smarter through the fight.

A new paper describes a training pipeline that pits an attacker language model against a defender model, updating each one in alternating rounds. The underlying reinforcement learning algorithm is GRPO, but scoring the fight is the harder problem: the system grades outputs with several large language model judges, checking things like attack success and defender helpfulness separately. To stop any one judge from dominating training, the researchers blend those reward channels with a separate technique called GDPO. Training starts with attacker-only single-turn and multi-turn practice before moving into full co-training between attacker and defender.

Jailbreak defenses usually rot the moment a new attack style shows up, so a pipeline that keeps retraining both sides against each other is a sturdier bet than patching against last month's prompts. The researchers' own ablations show which pieces of this setup actually move the safety-versus-utility needle, which is more rigor than most red-teaming claims get.

The honest catch: the paper admits GRPO tends to collapse attacker diversity over training, so a defender that beats its own sparring partner may still be unprepared for attacks nobody trained it against.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →