AI/ ai alignment · llm evaluation · ai safety · benchmarks

LLMs Nail Professional Ethics Calls, Flunk the Reasoning Behind Them

A new benchmark shows top LLMs mostly make the right ethics call in medicine, law, finance and security, but often cite the wrong reason for it.

A new benchmark says large language models mostly know the ethical rule, just not why it applies.

Researchers built ContextAdapt, a framework testing whether LLMs adapt values like honesty, autonomy, and confidentiality to the professional setting - medicine, law, finance, and national security - without drifting when the underlying rule hasn't actually changed. They built scenarios from real professional and regulatory documents, then scored 12 LLMs on both the action recommended and the justification given for it. Across the main test, models picked an appropriate action 95.6% of the time on average. But giving the correct, domain-specific reason for that action ranged wildly by model, from just 25.6% to 76.9%.

The more telling result is what didn't move the needle and what did. Explicitly naming the professional domain or changing the model's assigned role had little effect on behavior. Raising the perceived stakes did - in some cases models changed their answer even though the actual professional obligation hadn't changed at all, with severity apparently acting as a trigger for disclosing things that should have stayed confidential.

That's the gap that matters for anyone deploying these models in regulated work: a high pass rate on "did it do the right thing" can mask a model that got there for the wrong reason, and may fold under pressure the moment a scenario is dressed up to sound more serious.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →