Tell a reasoning model to act like a kindergartner, and it will happily keep solving calculus behind the cute voice.
Researchers built RoleCapBench, a benchmark testing three open-weight reasoning models across six educational roles and four assessment levels, from elementary school through A-level. They found that when models were prompted to adopt a lower-ability persona, they produced convincing in-role chatter - role-voice scores of 1.218 to 1.389 on the paper's scale - while still answering above-role questions correctly 81.1 percent to 89.8 percent of the time. That gap held even when prompts explicitly told the model to match the role's capability level. The team calls this "role-capability leakage," and it proved stubborn across prompting conditions.
This matters for anyone building role-play prompts into products: tutoring apps, kid-safe assistants, or personas meant to sound less expert than the system actually is. A model that talks like a six-year-old but reasons like a specialist isn't just quirky - it can undermine any product or safety design that assumes the persona reflects the actual capability ceiling. The researchers' proposed fix, called Injection, pairs explicit capability guidelines with a prefilled response prefix, cutting above-role accuracy by up to 0.562 while dropping in-role accuracy by less than 0.058 in most models.
Acting dumb, it turns out, is a costume change for these models, not a personality transplant - the calculus keeps happening backstage, no matter how convincing the kindergarten voice sounds.