[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-study-traces-ai-triage-test-failures-to-multiple-choice-format":10,"sections":35},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":24,"tags":25,"sources":30,"feedback":34,"feedback_at":22,"cost_usd":34,"total_tokens":34},10123,"study-traces-ai-triage-test-failures-to-multiple-choice-format","Study Traces AI Triage Test Failures to Multiple Choice Format","New interpretability research shows large language models understand emergency severity fine but stumble when forced to pick from multiple-choice options.","A new interpretability study suggests AI models aren't necessarily bad at medical triage - they're bad at multiple-choice tests.\n\nResearchers probed Gemma 3 4B and 12B instruction-tuned models, plus Qwen3-8B, using sparse-autoencoder features to trace how each model handles clinician-written emergency triage vignettes. The models correctly register emergency-severity signals in the case narrative itself, with decodability scores (ROC-AUC) between 0.95 and 1.00, regardless of whether the test is formatted as multiple-choice or free text. That signal weakens sharply at the exact moment the model has to pick a lettered answer. In the Gemma models, features tied to the multiple-choice format itself, not the medical content, accounted for over 91% of what drove the final answer, while the medical features contributed essentially nothing to that choice.\n\nThat is an important wrinkle for the growing pile of benchmarks claiming LLMs systematically under-triage patients. If the breakdown happens at answer selection rather than clinical understanding, then a model's multiple-choice score may say more about test design than about what it knows about emergencies. The researchers also found that whether the multiple-choice format helps or hurts varies by model, and the cases that flip between formats usually differ by just one severity tier, which rules out simple positional bias since the result holds up under option-order shuffles.\n\nWorth remembering next time a benchmark headline says a chatbot fails at triage: the test format might be the thing failing, not the model's medicine.","[\"llm evaluation\",\"ai interpretability\",\"healthcare ai\",\"benchmarking\"]","2026-10-05T04:00:00.000Z","2026-10-06T00:15:50.943Z","2026-10-06T00:15:56.841Z","published",null,[],"ai",[26,27,28,29],"llm evaluation","ai interpretability","healthcare ai","benchmarking",[31],{"name":32,"url":33},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2605.29889",0,{"sections":36},[37,41,45,50,55,60,64,69,74,79,84,89,94,99],{"name":38,"slug":24,"count":39,"latest_published_at":40},"AI",6317,"2026-10-05T09:51:57.000Z",{"name":42,"slug":43,"count":44,"latest_published_at":18},"Security","security",871,{"name":46,"slug":47,"count":48,"latest_published_at":49},"Policy","policy",446,"2026-10-05T10:25:00.000Z",{"name":51,"slug":52,"count":53,"latest_published_at":54},"Deals","deals",340,"2026-10-05T09:18:03.000Z",{"name":56,"slug":57,"count":58,"latest_published_at":59},"Hardware","hardware",205,"2026-10-05T10:58:22.000Z",{"name":61,"slug":62,"count":63,"latest_published_at":18},"Science","science",179,{"name":65,"slug":66,"count":67,"latest_published_at":68},"Consumer Tech","consumer-tech",160,"2026-10-05T10:23:15.000Z",{"name":70,"slug":71,"count":72,"latest_published_at":73},"Dev Tools","dev-tools",99,"2026-10-05T10:47:06.000Z",{"name":75,"slug":76,"count":77,"latest_published_at":78},"Software","software",97,"2026-10-04T10:00:00.000Z",{"name":80,"slug":81,"count":82,"latest_published_at":83},"Startups","startups",93,"2026-10-05T11:13:51.000Z",{"name":85,"slug":86,"count":87,"latest_published_at":88},"Gaming","gaming",53,"2026-10-02T02:50:39.000Z",{"name":90,"slug":91,"count":92,"latest_published_at":93},"General","general",51,"2026-10-05T02:35:01.000Z",{"name":95,"slug":96,"count":97,"latest_published_at":98},"Reviews","reviews",32,"2026-10-02T18:00:00.000Z",{"name":100,"slug":101,"count":102,"latest_published_at":103},"How-To","how-to",8,"2026-10-05T09:00:00.000Z"]