Language Models Can Tell When They're Being Tested
New research finds AI models internally register when they are being evaluated, and steering that internal signal changes what they admit aloud.
The Revision stories tagged llm-evaluation — 34 stories on the wire, decoded into one clear voice.
New research finds AI models internally register when they are being evaluated, and steering that internal signal changes what they admit aloud.
A new benchmark shows top AI coding agents lose accuracy on rewritten but equivalent code, and no model is robust across every tool.
A new benchmark testing whether AI assistants can handle weeks-long, ever-changing real-life tasks found seven top models all score poorly.
A new six-stage workflow tries to separate real psychological findings from measurement artifacts in large language model research.
A new review argues AI psychology research needs the same rigor as human studies, warning that results about AI personalities may be statistical noise.
A new study finds that self-consistency voting lowers accuracy on most hard GPQA problems for small LLMs, and simple confidence fixes don't help either.
Researchers ran a Milgram-style obedience test on 42 AI models and found compliance ranging from 0 to 100 percent, with tool calls cutting it sharply.
A new benchmark finds LLM data-analysis tools often misread spreadsheet structure, and that fixing this improves the accuracy of their answers.
A study of a 12,000-item Ukrainian judges exam shows top AI models can guess right answers from the options alone, exposing a benchmark design flaw.
A five-review study finds batching screening decisions changes LLM behavior more than prevalence metadata does, raising questions about hidden costs.
An audit of 37 AI models as survey stand-ins finds they skew positive, fabricate effects, and lose predictive accuracy versus real human data.
A benchmark shows mid-sized AI reasoning models reach right answers via flawed logic 28 percent of the time, exposing what answer-only grading misses.
A new benchmark called ALPS finds today's best AI models solve almost none of thousands of provably verifiable math construction problems.
LongDocBench finds AI document parsers ace single pages but struggle to rebuild tables of contents and figure-caption links in long documents.
A new paper finds that popular pass@k coding-agent benchmarks conflate test-suite size with independent attempts, inflating scores by up to 0.97.
A new benchmark shows even top language models can pinpoint what broke in a multi-agent system's execution only about a quarter of the time.
New research shows optimizing large language models for popular coding benchmarks does not reliably transfer to broader coding tasks.
A new benchmark finds four open-weight models refuse harmful Somali prompts far less often than identical English ones, often failing incoherently instead.
A new bilingual dataset grades legal RAG systems claim by claim, and even top performers stumble on both retrieval and generation.
A nine-agent framework verifies each claim in financial AI answers separately, raising accuracy and abstaining when evidence falls short.
A framework called optstop halts LLM benchmark tests once results are statistically certain, cutting trials up to 97 percent with the same conclusions.
A new benchmark finds frontier AI models still shift their answers toward planted numbers, even when they ace the same questions without hints.
A new benchmark compares four AI agent memory systems on multi-session tasks and finds no single approach wins across the board.
A new benchmark-free scoring method judges AI 'learning harnesses' by how closely they converge toward a stronger teacher model over time.
New research finds no single signal reliably predicts when an LLM update silently breaks a previously correct answer.
A new metric shows AI coding agents can succeed at the same rate while behaving wildly differently task to task, exposing a reliability gap benchmarks miss.
New research finds AI agents often keep a safety rule's wording after memory compaction while losing its effect, making simple presence checks unreliable.
A new benchmark finds emotional phrasing alone, with no numbers changed, cuts AI reasoning accuracy by 2 to 10 percentage points.
A new study maps scoring bias in AI evaluators to specific activation patterns, enabling bias prediction and even correction before a judgment is made.
A new study finds LLM-as-judge scores shift with evaluator upgrades, raising questions about whether benchmark results mean what researchers think they mean.
Researchers mapped 452 public LLM benchmarks onto 41 work activities and 38 banking domains, then used a weighted Elo system to score models for finance work.
A new four-stage diagnostic finds that top AI models can sense direction in unfamiliar physics but miscalculate ratios — and rarely catch their own mistakes.
A new empirical survey of 11 LLM evaluation setups confirms that accuracy, diversity, and consistency are fundamentally in tension with each other.
Researchers found that standard multi-model evaluation panels have a fatal flaw — one biased judge can corrupt the whole score, no matter how large the panel.