AI/ ai · clinical-ai · me-cfs · research

AI Agent Tries to Referee Conflicting ME/CFS Medical Guidelines

An unreviewed arXiv preprint says its ARCagent system hit 95.3% accuracy answering ME/CFS questions by explicitly modeling where medical guidelines disagree.

A new AI agent claims it can answer questions about ME/CFS correctly even when the official medical guidelines contradict each other.

The system, called ARCagent, is described in an arXiv preprint (arXiv:2609.36392) posted this week and has not been peer-reviewed. Its authors built a 1,706-chunk knowledge base pulled from 10 sources, tagged with a conflict registry that maps where competing ME/CFS diagnostic frameworks disagree on treatment. A retrieval pipeline then re-ranks evidence using those conflict signals instead of treating every source as equally reliable. The authors also built their own benchmark, scored by a separate LLM acting as judge rather than keyword matching, and report that ARCagent hit 95.3% accuracy, ahead of baseline LLMs, while arguing keyword-based scoring undercounts correct answers by 10.1 percentage points on average.

ME/CFS is a useful stress test for retrieval-augmented AI precisely because there is no single settled truth to retrieve. Diagnostic frameworks for the disease coexist and disagree, so an AI that just fetches the most relevant passage can confidently repeat one camp's position while ignoring the dispute entirely. The more interesting output here may be the conflict registry itself, a structured map of where medical authorities disagree, rather than the accuracy score attached to it.

All of this comes from one un-peer-reviewed preprint, with a benchmark the same team designed and an LLM judging the answers. The code is on GitHub, so the 95.3% figure is at least checkable, but treat it as a research claim, not a verdict.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →