OpenAI released GDPval, a benchmark built around a single idea: model performance should be measured in economic output, not test scores.
The company designed the evaluation to test its models against real tasks drawn from 44 occupations - work that people are paid to do. Instead of relying on academic tests or coding puzzles, GDPval maps AI capability directly to economically productive labor. The name makes the thesis explicit: GDP, not accuracy points, is the metric that matters here.
AI benchmarks have faced sustained criticism for being gameable and disconnected from practical value - a model can top a leaderboard while still failing at actual work. GDPval is OpenAI's answer to that critique: anchor the evaluation to real jobs, and a higher score should mean real economic utility. That argument is also convenient for a company that has spent considerable effort making the case to governments and investors that AI drives GDP growth.
A benchmark designed to show economic value, named after economic value, and built by the company whose models it evaluates is worth reading carefully.