A new academic review takes stock of how large language models are actually being used in medicine, and it lands more cautious than celebratory.
The paper, posted to arXiv, surveys recent talks and articles on applying LLMs to clinical research and practice. It contrasts these general-purpose models with older AI tools in medicine, which were typically built to handle one narrow task, such as reading a single type of scan. LLMs promise more flexibility, since they can work with varied, unstructured data rather than a single defined input. But the authors spend as much space on open questions around reliability, accuracy, transparency, and safety as they do on the potential upside.
That balance matters because hospitals are already piloting these tools quietly, and a model that drafts clinical notes or summarizes a patient history is only useful if it does not hallucinate a drug interaction along the way. The review reads less like a sales pitch and more like a checklist, which is a rarer thing in health-tech coverage than it should be.
Medicine has cycled through AI hype before, from narrow diagnostic algorithms to IBM's ill-fated Watson Health push, and each time the gap between demo and deployment turned out to be the real story. Expect the same gap to shape how LLMs fare in the exam room.