A new benchmark confirms that large language models still can't reliably count letters, find substrings, or spot palindromes - and it explains why in granular detail.
Researchers built SyntaxBench, a diagnostic framework that tests character-level reasoning instead of lumping it into one vague accuracy score. It runs five core tasks - character counting, letter containment, palindrome detection, edit distance, and longest-string selection - using matched pairs of English words and random strings of the same character length, plus a tougher substring-extraction task called index_to_span. The team tested eight open-weight models ranging from 2 billion to 32 billion parameters, across 11 different reasoning-mode setups and zero-, one-, and four-shot prompts. Instead of reporting a single score, they ran statistical tests - McNemar's test, Cohen's kappa, bootstrap confidence intervals - to separate real differences from noise.
The headline result is that tokenization, not reasoning ability, drives a lot of the failure. Random strings get split into roughly 1.9 characters per token versus 3.2 for English words, and counting accuracy drops as words get swallowed into fewer, bigger tokens - the model is working from a blurrier picture of the string than you'd assume. Explicit reasoning, or 'thinking mode', does not reliably help either: one 27-billion-parameter model actually got worse at spotting palindromes when allowed to reason it out, dropping from 95.2% to 88.6% accuracy at four-shot.
The substring-extraction task, index_to_span, remains basically unsolved at 6.75% best-case accuracy - so the next time a model miscounts letters in a word, the fix isn't more prompting, it's a different way of looking at text altogether.