AI/ ai · llm-benchmarks · deepseek · open-source

Open-Source Reasoning Model Benchmarks Don't Reproduce

A new study finds DeepSeek-R1-Distill and QwQ-32B benchmark scores swing wildly with small evaluation changes, undercutting claimed reasoning gains.

The benchmark scores that made DeepSeek's open reasoning models look so good may not hold up under scrutiny.

Researchers re-ran evaluations on the DeepSeek-R1-Distill series and the QwQ-32B model and found that scores fluctuated substantially depending on evaluation conditions. The same instability showed up in other open-source models fine-tuned on top of the R1-Distill base. That makes the performance gains those models claim on math, science, and coding benchmarks hard to reproduce reliably. The researchers ran their own empirical assessments of the R1-Distill models and are pushing for a more rigorous, standardized evaluation paradigm.

This matters because open-source reasoning models have leaned hard on leaderboard numbers to argue they can hang with closed frontier models. If those numbers shift based on how the test is run rather than what the model actually knows, every benchmark screenshot circulating online deserves a second look. It also puts developers choosing a model off published scores in a tougher spot.

Benchmark gaming gets talked about a lot. This is a quieter problem: the measuring stick itself may be too noisy to trust.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →