AI/ ai · nlp · healthcare · speech

LLMs Can Score Aphasic Speech, But Not Without Examples

A small-model study finds few-shot prompting closes most of the gap with trained human raters on aphasia discourse scoring, but precision remains a problem.

Small language models can score aphasic speech transcripts nearly as well as trained clinicians, but only when given examples to work from first.

Researchers tested four instruction-tuned models on Correct Information Unit (CIU) scoring, a clinical measure of communicative informativeness that counts what a speaker conveys, not just how they say it. The models processed 16 picture-description transcripts covering mild to severe aphasia. Zero-shot prompting failed across all four. With few-shot examples, three of them (Llama-3.1-8B, Qwen2.5-7B, and Mistral-7B) reached F1 scores between 0.776 and 0.817 against consensus human labels; Phi-3-mini was unreliable throughout. How the examples were selected, fixed global versus per-chunk local, made no significant difference to results.

CIU scoring currently requires trained speech-language pathologists to annotate transcripts word by word, a time cost that limits where aphasia assessment is realistically available. Automating even partial scoring could expand access in under-resourced clinics and telehealth settings where a specialist is not on hand. The obstacle is a precision problem: all three viable models over-classified tokens as CIUs, catching too much rather than too little, which in clinical practice could inflate a patient's apparent communicative ability.

Performance dropped most sharply on severe aphasia cases, which is exactly where faster automated tools would matter most. The researchers frame this as a promising human-in-the-loop result, but the best numbers come from the easier end of the severity spectrum, a gap the paper acknowledges without resolving.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →