A new paper tries to answer a basic but oddly unresolved question: when a clinical trial says AI was involved, what does that actually mean?
Researchers built a multidimensional framework for categorizing human-AI interactions in clinical trials, sorting them by what task the AI performed, how it relates to the humans involved, how the interaction is configured, and which human groups are doing the interacting. They tested it on 15 clinical trials pulled from an existing dataset, with two human reviewers and six large language model classifiers independently categorizing each one. The LLMs could handle much of the categorization work on their own. But when trial records were incomplete or vague, human judgment still won out.
That inconsistency is the real story here. A trial summary describing AI assistance can mean a model flags a result for a clinician to double-check, or it can mean the model makes the call with no human in the loop - two very different things lumped under the same label. A shared framework like this one gives researchers doing systematic reviews a common vocabulary instead of guessing from vague trial descriptions, which matters as regulators and journals try to compare AI trials against each other.
Fifteen trials is a thin test bed, and the people who built the framework are also the ones who validated it. Whether this categorization scheme holds up against the next few hundred AI clinical trials - rather than the handful chosen to prove it works - is the question worth watching.