AI/ ai · education · llm-evaluation · research

AI Essay Feedback Skips the Teaching Part

A new benchmark finds leading language models can mimic what teachers say about student writing but not when or why they say it.

Researchers have built a benchmark showing that AI models giving feedback on student essays hit the right talking points but miss the teaching behind them.

The study, released on arXiv, tested six large language models under three different prompting setups against real feedback from instructors across three university writing courses. The team adapted an existing framework from education researcher Susanne Narciss into seven categories of feedback focus, covering things like grammar, structure, and argument quality, then annotated both teacher and AI-generated comments to compare them directly. The resulting dataset, called FeedType, lets researchers measure not just what feedback says but where it's aimed and whether it changes as a student's work evolves.

The key finding: most models can produce comments that touch all seven feedback categories, but they don't weight them the way teachers do, and none of the six adjusted their feedback style based on draft stage or student skill level the way human instructors did. That distinction matters more than it sounds. A teacher naturally backs off line-edits on a rough first draft and leans into structural comments, saving grammar nitpicks for later revisions. An AI tool that can't make that call risks burying a struggling first-draft writer in copyedits instead of the bigger-picture help they actually need.

It's a useful corrective to the current wave of AI writing-feedback tools, which tend to be graded on whether their comments sound plausible rather than whether they teach anything. Coverage of the right topics isn't the same as knowing when to bring them up, and this benchmark is one of the first attempts to measure that gap directly rather than assume it away.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →