AI/ ai-training · synthetic-data · llm-benchmarks · reasoning-models

Researchers Teach AI to Write Its Own Harder Test Problems

A self-improving harness that writes its own harder reasoning problems already produces training data strong enough to rival frontier-model math benchmarks.

A new AI research harness teaches itself to write harder test questions, round after round.

Researchers built a system called task-harness co-evolution that goes a step beyond earlier approaches to synthetic training data. Older methods reused generated problems as seeds for new ones but left the problem-writing process itself untouched. This one evolves the harness too: it turns a solver's mid-generation failures into reusable skills, then after each batch revises its own skills, prompts, and workflows, keeping only changes that produce genuinely harder, valid problems without a runaway increase in cost. Over fourteen rounds spanning math, coding, and science, average solver accuracy dropped from a perfect 100% to 54.8%, with the model's weights and its answer-checking criteria held fixed the whole time.

That accuracy drop is the real story: the system is manufacturing problems that get harder on their own, without a human writing trickier questions by hand. That matters because frontier AI labs are increasingly bottlenecked on hard, verifiable training data, not raw compute. The paper's headline result backs this up: a 27 billion parameter model fine-tuned on just 10,000 of these synthesized math problems scored 62.5% mean accuracy on the APEX benchmark, which the authors call competitive with selected frontier-model references.

Selected is the word to watch there - it is not a claim of beating every frontier model, just a chosen few comparison points, about as close as a methods paper gets to marketing copy.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →