A new framework grades how risky a piece of health information really is based on context, not just the words used.
Researchers built a framework for evaluating sensitivity in online medical conversations, like forum posts or chat-based consultations, that goes beyond simple entity labeling. It tracks four context signals drawn from clinical documentation standards: whether a condition is confirmed, suspected, negated, or hypothetical; who the symptom belongs to; whether a related test result is pending or returned; and how specific the disclosure is. The team built contrastive test cases - near-identical conversation snippets that alter just one of those signals at a time - and used them to compare large language models that see only a bare mention of a condition against models given the full surrounding context.
Privacy tools that flag sensitive health disclosures are only as good as their ability to separate a real diagnosis from a suspected, negated, or hypothetical one. Mislabeling a hypothetical or negated condition as confirmed could trigger unnecessary alerts or over-redaction, while missing a genuinely confirmed disclosure could let sensitive data slip through unflagged. The framework is built to measure both of those failure modes separately, rather than folding them into one accuracy score.
It is a reminder that most AI classification benchmarks test what was said, not what was meant - and in a medical chat log, that distinction is the whole point.