General-purpose language models just outscored astronomy-specialized ones on a real scientific-reasoning benchmark.
A new paper, "Rethinking Domain Specialization for Open-Ended Scientific Reasoning in Astronomy Language Models" (arXiv:2609.17644), tested general-purpose, multimodal, and astronomy-specialized language models against 300 free-response questions pulled from 2017-2026 astronomy Olympiad materials. The set splits into 204 text-only questions and 96 that require reading an image. The paper's authors judged answers with a judge-based correctness method plus several reference metrics, comparing both open-weight and API-served models. The strongest general-purpose models posted the highest correctness scores in the test, ahead of the models built specifically for astronomy.
That's a real problem for the pitch behind domain-specific fine-tuning: if a general model already reasons through Olympiad-level astronomy better than a specialized one, the case for building niche science models gets narrower. The paper's own breakdown adds a caveat, though: judge sensitivity, benchmark composition, and whether a question includes an image all moved the rankings, so no single leaderboard score captures the full picture.
Worth remembering next time a startup pitches a "domain-specific" model as inherently better at science: on this benchmark, specialized lost to general, and the details of how you measure it mattered as much as which model you picked.