A new evidence-retrieval method lets AI models think several steps ahead before deciding which facts to trust, and it beats prior approaches on a standard benchmark by double digits.
Researchers built a system called Foresight-over-Graph (FoG) that helps large language models search knowledge graphs for question answering. Unlike existing methods that prune weak-looking evidence hop by hop as they go, FoG builds a question-relevant subgraph and uses feedback from deeper exploration to decide which paths were actually worth keeping. On the CWQ benchmark, a standard test for complex, multi-hop questions, FoG improved the Hit rate - how often the right answer shows up among the retrieved evidence - by 16.58%, while making fewer LLM calls and using fewer tokens than competing methods. The code is public on GitHub.
Knowledge graphs are one of the more credible pitches for keeping LLMs honest on fact-heavy questions, since a graph is structured and checkable in a way a model's internal weights are not. The problem with earlier graph-search methods was that greedy, hop-by-hop pruning threw away branches that only turned out to matter several steps later, with no way to recover them. Fixing that ordering problem, rather than adding more training data, is what moved the number.
A hit-rate gain on one benchmark is not proof that graph-grounded LLMs have solved hallucination - it is evidence that smarter evidence-retrieval helps on this particular, multi-hop-heavy test, and that is a narrower claim worth keeping straight.