BioASQ, the long-running biomedical AI benchmark, just wrapped its fourteenth round - and the field keeps growing, not shrinking.
The challenge, held alongside the CLEF 2026 evaluation forum, split into six tasks this year. Task 14b covered standard biomedical question answering, while Synergy14 tackled QA for still-developing biomedical topics. MultiClinSum-2 tested multilingual clinical summarization, BioNNE-R focused on extracting relations between nested named entities in Russian and English, ELCardioCC handled clinical coding in cardiology, and GutBrainIE targeted gut-brain interaction extraction. In total, 87 teams entered and submitted more than 1,000 runs, with several reaching what organizers call competitive performance.
That spread of tasks is the real story. A benchmark that started out asking whether a model could answer a biomedical question now has dedicated tracks for cardiology coding and gut-brain biology - a sign that biomedical NLP research has matured past generic QA and into clinical specialties with their own jargon, formats, and failure modes. Eighty-seven teams and over 1,000 runs also suggests this corner of AI research is crowded enough that incremental gains, not breakthroughs, are the norm.
None of that is visible in the numbers alone. The organizers don't name which teams or methods actually won, so competitive performance is doing a lot of unverified work here. Until a public leaderboard or follow-up paper names names, treat this as a census of effort, not a verdict on progress.