AI/ ai · gpu-kernels · benchmarks · llm-agents

D2K-Bench Tests How Well AI Agents Write GPU Kernels

A new diagnostic benchmark shows AI coding agents still need expert hand-holding to write GPU kernels that approach expert-level performance.

A new benchmark finds that AI coding agents need detailed human blueprints before they get anywhere close to expert-level GPU kernel performance.

Researchers built D2K-Bench, a set of 26 tasks and 85 workloads that tests whether large language model agents can turn expert design guidance into fast GPU kernels. Each task ran twice: once with written guidance covering algorithm choice, data flow, and low-level optimization tricks, and once without it, using identical task descriptions, tools, hardware, and a 350-turn budget on NVIDIA B200 GPUs. Across five models and 130 model-task pairs, guidance pushed correctness from 93.1% to 98.5% and raised the benchmark's Performance Score from 1.46 to 1.95. For the three models that solved every task correctly in both runs, GPT-6-Astra, Claude-Opus-4.8, and GPT-5.6-Sol, geometric mean speedup rose from 1.69x to 2.49x, a 47% improvement, not the near-doubling it might look like at first glance.

That gap matters because it separates two different skills we keep lumping together: writing code that runs, and writing code that performs. Hand an agent the design blueprint and it executes competently. Ask it to discover that blueprint on its own, and the performance gap widens substantially.

It is a useful corrective for anyone reading "AI writes code now" as "AI writes code like a performance engineer" -- this benchmark is built specifically to measure the difference.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →