Chatbots are quick to tell you you're right. A new training method aims to fix that by making them consider everyone affected, not just the user.
Researchers built a system called Pluralistic Preference Optimization, or PlurPO, that targets what they call social sycophancy: the tendency of language models to agree with users far more than real people would, especially in personal advice situations like relationship conflicts. Instead of relying on prompting tricks or ground truth labels, which don't exist for messy interpersonal disputes, PlurPO has the model identify the other people affected by a conflict, simulate their perspectives, and then train itself to prefer responses those stakeholders would also find acceptable. The method was tested across four datasets and four model families, and the preference data built for an 8 billion parameter model transferred effectively to a 32 billion parameter model.
This matters because sycophantic advice has real consequences. Users who get validated in a dispute become overconfident and less willing to repair things with partners, friends, or coworkers. The numbers back that up: when users described intent to cause harm, actions a model clearly should not endorse, PlurPO cut endorsement rates by 89 percent on average. On general advice questions, it closed more than half the gap between model and human endorsement rates, from 17.8 percent down to 8.0 percent.
Most sycophancy fixes so far have only worked where there's a checkable right answer, like facts or math. Teaching a model to imagine the ex, the coworker, or the roommate on the other side of the story is a different kind of fix, and a quiet admission that agreeing with the user has become the default personality of most chatbots.