Chatbots are becoming parenting coaches, and a new benchmark suggests they're better at looking helpful than being consistently helpful.
Researchers built a multi-dimensional rubric with input from parenting experts and used it to grade 15 large language models on 100 parenting scenarios, in both English and Chinese, relying on an LLM-as-a-judge method to score the responses. The study found that a model's overall average can mask weaknesses tied to specific rubric items - a model can look solid in aggregate while quietly failing particular kinds of advice. It also found that models implicitly steer parents toward different parenting styles without disclosing that they're doing so, and that the language a parent writes in changes the advice they receive.
That matters because parenting isn't a domain where 'mostly fine' cuts it. A single leaderboard score won't tell a developer or a parent whether a model is bad at exactly the scenario they're facing, like discipline or safety questions, and the language-dependent results suggest non-English speaking parents could get systematically different guidance from the same app. The researchers frame this as an argument for auditable evaluation output, not just a summary score.
It's the same lesson AI evaluation keeps relearning elsewhere: a tidy benchmark number is easy to market and easy to misread, especially in a domain where the wrong answer isn't a bug report, it's advice a parent might actually follow.