GPT-5 can grade preschool teachers about as well as trained observers, but only on the easy parts.
Researchers fed GPT-5 transcripts from 87 video-recorded classroom observations across 38 classrooms in 30 Hong Kong kindergartens. The model applied the Classroom Assessment Scoring System, the standard framework used to judge how teachers interact with young children, and its ratings were compared against scores from trained human raters. The AI matched humans closely on Emotional Support, especially the Quality of Feedback dimension, which tracks whether teachers build on what children say. It diverged more on Classroom Organization and Instructional Support, the domains that depend on procedure and context rather than verbal exchange.
Classroom observation is expensive precisely because it requires trained humans watching video and applying a rubric consistently, which is why schools do it rarely despite its value for teacher development. An AI shortcut is tempting, but this study is really a map of where that shortcut breaks: language-heavy, feedback-focused moments translate well to a transcript, while spatial and procedural judgment calls do not. That mirrors a pattern seen elsewhere in AI evaluation tools, from essay grading to code review, where models handle content quality better than they handle process and context.
The researchers themselves suggest AI belongs in the pile of tools teachers use for self-reflection, not the file HR pulls at review time.