AI/ ai · multi-agent-systems · reinforcement-learning · research

New Method Lets AI Agents Cooperate Without Seeing Rewards

A new training method lets AI agents cooperate in resource-sharing games by inferring peers' rewards from behavior instead of peeking at private scores.

A new reinforcement-learning technique teaches AI agents to cooperate without ever seeing what their fellow agents are rewarded for.

The method, called self-referenced social preferences, has each agent learn a model of its own reward and apply that model to what it observes other agents doing, instead of requiring direct access to those agents' actual reward signals. Researchers tested it in three simulated social dilemmas: Escape Room, which rewards volunteering; Clean Up, a public-goods contribution task; and Commons Harvest, a shared-resource game that punishes overconsumption. Agents trained this way learned to cooperate in cases where agents with no social preference at all, learning independently, failed to cooperate. In several tests, these agents ended up splitting the rewards more evenly than agents given direct access to everyone's true reward signals.

Most prior work on cooperative multi-agent AI assumes each agent can peek at its peers' private reward signals, a convenient research shortcut that doesn't match how cooperation actually works. This approach drops that assumption and relies only on observed behavior, which matters for any future multi-agent system, think trading bots, warehouse robots, or AI assistants coordinating, that won't have transparency into each other's incentives.

It's a simulation with three toy games, not a commons of self-driving cars, so treat the equity numbers as a promising lab result rather than a settled finding.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →