A 9-billion-parameter AI model just out-diagnosed GPT-5.5 on rare diseases by leaning on a disease knowledge graph instead of brute scale.
The work, described in an unpublished preprint posted to arXiv (2609.35549), is called RareDx. Its authors built RareDx-Harness, a benchmark that folds messy, incomplete patient records into a single ranked-diagnosis task, then trained the model to reason over phenotypes, genes, and diseases using a shared disease knowledge graph rather than open-ended web retrieval. The training method, RareDx-KGPO, scores answers by how close they land to the correct disease inside that graph and caps the length of a model's guess list, so it cannot rack up partial credit by listing plausible-sounding diagnoses. Across eight benchmarks, the 9-billion-parameter system landed the correct diagnosis in its top 10 guesses 38.34% of the time on average, 1.6 points ahead of GPT-5.5 under the same archived test protocol; a larger 27-billion-parameter version scored 23.53%, 36.56%, and 40.76% at top-1, top-5, and top-10.
That gap matters because rare-disease diagnosis is exactly where general-purpose chatbots struggle: the conditions are individually rare, sparsely documented, and easy for a model to either miss or fabricate a plausible-but-wrong name for. A relatively compact model beating a frontier system suggests that structure, in this case a curated graph of phenotypes, genes, and diseases, can substitute for raw scale. That is a cheaper, more auditable approach for hospitals or health systems that cannot run frontier-scale infrastructure on every case.
The paper's own ablations add a useful caveat: bolting on retrieval alone did not reliably help, and the routing and output constraints were what actually drove the gain, not extra data access. That fits a familiar pattern in diagnostic AI, where benchmark leaderboards move years ahead of anything reaching a doctor's screen. RareDx was tested against archived records and rival algorithms, not against a clinician in a real workup, so the real test is whether a graph-grounded model holds up against data messier than any benchmark.