A new training method lets small AI agents learn to check their own homework instead of relying on expensive, human-written answer keys.
The method, called GraphCert, targets graph agents - language models that explore and reason over knowledge graphs through multi-step tool calls - and tackles the field's usual bottleneck: training data. Normally that means paying people to write question-answer pairs and reasoning traces by hand, or sending proprietary graph data to an outside AI provider to generate them instead, both slow, costly, or risky for sensitive information. GraphCert skips both: one internal component generates its own graph-grounded questions and supporting evidence, checks that evidence by actually running it and discarding anything that fails, then turns what survives into a scoring rubric. That rubric rewards a second component for citing the right evidence and reaching the right answer during training, rather than grading answers alone.
On GRBENCH, a benchmark spanning five graph domains, the resulting compact agent beat substantially larger LLM agents and rival post-training methods. It also held up when moved to graph domains it had not trained on, suggesting it learned transferable reasoning habits rather than memorizing quiz patterns for one dataset. That combination - smaller models beating bigger ones without shipping data to a third-party API - matters for any company sitting on a proprietary knowledge graph it would rather not hand to an external model provider just to generate training data.
Worth noting: this is a preprint, the code is only promised and not yet released, and "substantially larger" is relative - GRBENCH is a research benchmark, not a messy production graph.