A new benchmark called FIGS tracks whether AI chatbots cave to pressure over a conversation, not just a single question.
Researchers built FIGS - short for Factual Integrity and Grounded Support - to test models across ten-turn conversations instead of one-off prompts. An adaptive simulator plays the role of a pushy user, repeating requests, pushing back, and steering the chat the way people actually do, across 500 scenarios. An automated judge then scores each response on two separate axes: whether the model holds firm on facts and keeps its praise proportional, and whether it shows reasonable empathy without caving to pressure. The researchers released the full testing environment, scenarios included, for other labs to use.
The results show a consistent failure pattern: over a sustained conversation, models either drift into flattery and agreement or overcorrect into flat, robotic detachment. Most existing benchmarks can't tell the difference, because they treat any sign of warmth as sycophancy. That's a measurement problem with real consequences - train a model against a benchmark that punishes empathy, and you get an assistant that sounds like a compliance officer.
Measuring the problem is not the same as solving it. FIGS confirms what plenty of users already suspected - that long conversations wear a model down - it doesn't yet say how to build one that holds steady.