AI models that read code start to fail at exactly the moment reading code gets hard.
A new benchmark tests two code-focused models, DeepSeek-Coder-V2 and Llama, on how well they understand Python functions as those functions get structurally messier. Researchers scored complexity using four established metrics - cyclomatic complexity, nesting depth, branching factor, and Halstead volume - then split 300 functions into low-, medium-, and high-complexity bands. On an automatic input-output prediction task, DeepSeek-Coder-V2 hit 78.33% accuracy overall and Llama hit 70.33%. Broken out by complexity, DeepSeek-Coder-V2 fell from 93.52% on simple functions to 52.78% on the hardest ones, and Llama dropped from 87.04% to 47.22%. A smaller, manually graded check of 60 functions showed the same slide: DeepSeek-Coder-V2 went from 100% to 75%, Llama from 90% to 60%.
That gap matters more than the overall accuracy figure. Code review, debugging, and refactoring in real codebases rarely involve the tidy, low-complexity functions these models handle well - they involve the tangled ones. A model that's reliable on toy examples but close to a coin flip on deeply nested, branch-heavy code is not one you'd want reviewing a pull request unsupervised.
Vendors keep publishing single accuracy numbers because single numbers make good marketing copy. This paper is a reminder to ask what the hard tail looks like before trusting one.