AI/ llm evaluation · ai interpretability · healthcare ai · benchmarking

Study Traces AI Triage Test Failures to Multiple Choice Format

New interpretability research shows large language models understand emergency severity fine but stumble when forced to pick from multiple-choice options.

A new interpretability study suggests AI models aren't necessarily bad at medical triage - they're bad at multiple-choice tests.

Researchers probed Gemma 3 4B and 12B instruction-tuned models, plus Qwen3-8B, using sparse-autoencoder features to trace how each model handles clinician-written emergency triage vignettes. The models correctly register emergency-severity signals in the case narrative itself, with decodability scores (ROC-AUC) between 0.95 and 1.00, regardless of whether the test is formatted as multiple-choice or free text. That signal weakens sharply at the exact moment the model has to pick a lettered answer. In the Gemma models, features tied to the multiple-choice format itself, not the medical content, accounted for over 91% of what drove the final answer, while the medical features contributed essentially nothing to that choice.

That is an important wrinkle for the growing pile of benchmarks claiming LLMs systematically under-triage patients. If the breakdown happens at answer selection rather than clinical understanding, then a model's multiple-choice score may say more about test design than about what it knows about emergencies. The researchers also found that whether the multiple-choice format helps or hurts varies by model, and the cases that flip between formats usually differ by just one severity tier, which rules out simple positional bias since the result holds up under option-order shuffles.

Worth remembering next time a benchmark headline says a chatbot fails at triage: the test format might be the thing failing, not the model's medicine.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →