A new benchmark finds that AI agents writing GPU kernels post big speedups in testing and far smaller ones in the real inference systems those kernels actually run in.
Researchers at Snowflake AI Research built FastKernels, a set of 384 kernel-writing tasks spanning 47 model architectures, built to mirror the modules inside production frameworks rather than isolated sandboxes. The tasks cover enough ground to reimplement 94.6% of HuggingFace Transformers architectures with outputs matching the native code. The team ran five AI agents, including Claude Code and KDA, through 6,900 agent-hours of kernel generation, then scored the results both in isolation and end to end inside real models using a new metric called MacroEval. Kernel-level speedups of 1.6x to 6.6x shrank to at most 1.25x once the kernels were plugged into full models, and only 20% of the kernel sets that won in isolated testing ran correctly as-is in production.
The gap matters because most GPU kernel benchmarks today score kernels alone, on synthetic inputs, which is exactly the setup FastKernels shows is misleading. Claude Code matched or beat KDA on every isolated kernel metric, yet KDA scored three times higher once both were evaluated inside actual inference pipelines. That's a full reversal that would be invisible to anyone benchmarking kernels the old way.
It's the same lesson AI benchmarking keeps relearning: optimize for the test and you get agents that are excellent at the test, not at the job.