AI/ ai benchmarks · data science · llm evaluation · code translation

New Benchmark Finds AI Models Struggle to Translate Data Science Code

Claude Opus 4.6 topped a new ORCA benchmark but still got only 33.67% of full-project code translations right, and 56.92% on simpler ORCA-MAIN tasks.

A new benchmark shows that even top AI models struggle to translate data science code from one library to another without breaking it.

Researchers built ORCA, a benchmark for what they call Data Science Code Translation: converting code between data science libraries while preserving its exact behavior. It has two parts. ORCA-MAIN covers 1,600 narrow tasks across data querying, data manipulation, and deep learning. ORCA-PROJECT covers 200 full, real-world data science projects across seven task types, with every translation checked against reference solutions and automated test cases.

Claude Opus 4.6 topped the leaderboard, but its scores expose a wide gap between flashy code generation and the messier task of migrating existing code: 56.92% success on the narrow ORCA-MAIN tasks, and just 33.67% on full ORCA-PROJECT translations. Models did notably better when the source code spelled out each operation explicitly rather than relying on implicit shorthand, a pattern the researchers exploited by having models summarize a snippet's intent before translating it, which lifted scores by about five points on both benchmarks.

Even with that intent-summarizing boost, getting a full project translation right remains roughly a 1-in-3 shot, closer to rolling a die than flipping a coin.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →