Almost none of last year's top AI papers reported how much carbon they burned to train their models.
Researchers ran an automated review of all 5,285 papers accepted to NeurIPS 2025, the field's largest conference, and found that environmental-impact reporting is essentially absent. In response, they built a set of standardized sustainability metrics for measuring training efficiency, along with rough heuristics for estimating the carbon cost of running a model after it is deployed. Those metrics ship in a new open-source tool called carbonbenchmark, meant to plug into existing training pipelines. The team also proposes a framework called Smallest Model that Achieves the Job, or SMAJ, which asks researchers to justify compute spend relative to accuracy gains instead of chasing state-of-the-art scores by default.
AI labs have spent years racing for leaderboard wins while treating energy use as someone else's problem. A standardized way to measure and report carbon costs gives conference reviewers, funders, and companies a concrete number to weigh against marginal accuracy bumps, instead of taking efficiency claims on faith.
Conferences already require ethics statements; whether NeurIPS or its peers make a carbon line item mandatory is the real test of whether this becomes more than a good idea.