AI/ ai · llm-evaluation · benchmarks · code-generation

New Benchmark Tests Whether AI Just Memorized the Algorithm

A new benchmark, AlgoREval, finds LLMs recall algorithms unevenly across languages and formats, questioning trust in AI-written code.

A new benchmark asks a blunt question: when an AI writes textbook code, is it reasoning or just quoting from memory?

Researchers built AlgoREval, a set of 599 problems covering 77 classical algorithms across 14 domains, 7 programming languages, and 4 ways of representing graph inputs. They tested 15 open models, ranging from 7 billion to 34 billion parameters, with no extra prompting, specifically checking whether models could reproduce known algorithms rather than invent new ones. Accuracy swung a lot depending on the programming language and how graph inputs were formatted, even for algorithms that are extremely well documented online. The team also found that feeding models retrieved code snippets or structured hints boosted scores on harder algorithms, and that standard fine-tuning spread gains across languages while reinforcement-style training (GRPO) produced bigger but narrower per-language improvements.

The paper's real contribution is a framing shift: it argues that writing code and recalling code are different skills that current benchmarks conflate. That matters because a model acing algorithm tests might just be pattern-matching its training data, not reasoning about the problem. For anyone shipping AI-generated algorithmic code into production, the finding has teeth - if retrieval wobbles this much by language or input format, the same model can look sharp in Python and shaky in Rust on the identical algorithm.

It is a quieter, more useful kind of AI paper than most - no leaderboard chest-thumping, just a reminder that fluency and understanding are not the same thing, and that gap is exactly where production bugs live.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →