AI/ text-to-3d · ai-benchmarks · evaluation-methods · generative-ai

Text to 3D Leaderboards Shift When You Just Change the Camera

A new audit finds that tweaking render and caption settings, not the 3D models themselves, flips which text-to-3D generator looks best on 18 of 19 evaluators.

Benchmark rankings for text-to-3D generators can flip even when the 3D scenes never change.

A new arXiv audit tested 300 frozen scenes from six text-to-3D generators against 19 alignment evaluators and one perceptual-quality control. The researchers kept the generated scenes fixed and only varied eight render and caption factors, like camera angle and wording, used to measure them. For 17 of 19 evaluators, the variance caused by these configuration choices was bigger than the variance between different generators. Scores moved more than rankings did, but 18 of 19 evaluators still changed which generator they crowned the winner under some configuration.

This matters because text-to-3D leaderboards get cited as evidence one model is better than another, often to justify funding, adoption, or press coverage. If the evaluation pipeline, not the model, is driving who wins, those claims are shakier than they look. The paper is careful to note that none of the observed winner-swaps survive rigorous statistical testing across the full range of settings, so this isn't proof any specific ranking is wrong, but it is proof the measurement process itself is fragile.

The researchers also found that even their best evaluator, including a control with no text prompt at all, correctly detected scrambled layouts only 67% of the time. That's a reminder that current automated judges for 3D generation are still closer to rough guesses than to reliable referees, and that the field's benchmarks deserve the same scrutiny as the models they rank.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →