AI/ ai · healthcare · gpt-4 · quality-assurance

Study Finds GPT-4 Wildly Overflags ER Revisit Cases

In a small study, GPT-4 flagged 94 percent of ER revisit cases for review, far more than clinicians, prompting researchers to test a more targeted algorithm.

A new study finds GPT-4 flags almost every emergency department revisit as worth a second look, while human reviewers agree on far fewer.

Researchers at a multihospital health system retrospectively reviewed 99 diagnosis pairs from ED revisits occurring 1-14 days after an initial visit. Given only the primary diagnosis from each visit, they asked 2-3 clinicians and GPT-4 to judge whether each pair warranted further quality review. GPT-4 flagged 94% of pairs as worth following up, 4.4 to 13.3 times more often than clinicians did, despite minimal prompt engineering. The team then built a knowledge-graph algorithm, called KGA, that used an LLM to populate structured relationships between diagnoses; it hit 83-100% positive predictive value for matching at least one clinician's call for review.

ED quality reviews are usually restricted to a narrow window, often 48-72 hours, specifically to keep the chart-review workload manageable, which means slower-developing problems can go unexamined. This study suggests that simply pointing a general-purpose model like GPT-4 at the problem produces so many false positives it would swamp reviewers with busywork, but a more structured, knowledge-graph-based approach looks like a workable middle ground.

Call it a reminder that reading a chart is not the same as exercising clinical judgment, and that hospitals shopping for AI screening tools should ask for positive predictive value, not a demo.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →