A new browser tool pings you mid-conversation when your AI assistant might be making things up, flattering you, or sounding more confident than it should.
The tool, called Safety Nudges, watches chatbot conversations in the browser and surfaces lightweight flags when it detects patterns like hallucination, sycophancy, overconfidence, or anthropomorphism. Researchers tested it in a two-week field study with 45 people who chat with AI systems frequently, gathering interaction logs, surveys, and reactions to individual nudges. Participants called the tool useful, clear, and not too intrusive, and nearly all said it made them more aware of how AI systems can go wrong. But that awareness mostly stayed in participants' heads: the study found it did not reliably translate into changed behavior, like double-checking a claim or rephrasing a leading question.
That gap is the real finding here. It is one thing to tell users a chatbot might be sycophantic; it is another to get them to act on that in the moment, especially when the whole point of a chatbot is to move fast and trust the output. The result lines up with older research on nutrition labels and privacy nudges, where informing people rarely changes what they do next. For an industry that treats "add a disclaimer" as a safety strategy, this is a useful data point that awareness and behavior are not the same problem.
In other words: the warning light works. Whether anyone still pulls over is a separate question.