AI/ llm evaluation · recommendation systems · ai benchmarks · user modeling

New Benchmark Says AI Models Misread Your Interests

A new benchmark called GISTBench shows that even top models like GPT-5 and Claude 4.6 struggle to accurately infer user interests from engagement data.

A new benchmark says today's best AI models are bad at figuring out what you actually like.

Researchers built GISTBench, a benchmark that tests whether large language models can correctly infer a user's interests from their interaction history, rather than just predicting the next item they will click. Using a synthetic dataset built from real engagement data on a global short-form video platform, the team scored eleven models - eight open-weight systems ranging from 7B to 235B parameters, plus GPT-5, Claude 4.6, and Gemini 3.5 Flash - on two new metrics: Interest Groundedness, which checks whether predicted interests are hallucinated or actually backed by evidence, and Interest Specificity, which checks whether a predicted profile is distinctive rather than generic. The team validated the synthetic dataset against real user surveys to confirm it reflected genuine behavior.

Recommendation systems increasingly lean on LLMs to summarize what a user wants rather than just rank what they might click next, and this benchmark finds that models, including the frontier ones, routinely miscount and misattribute engagement signals across different interaction types, like conflating a video someone skipped with one they watched twice. That is a basic accounting problem, not a subtle judgment failure, and it suggests the "AI understands you" pitch behind a lot of personalization products is still aspirational.

Bigger parameter counts have not solved an arithmetic problem: a model that can write a paragraph about your interests still might not be able to count how many times you actually watched something.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →