A new technique lets AI agents build lasting expertise on a document instead of searching it from scratch every time.
Researchers describe SourceLearn, a method that builds a persistent source model capturing how a document's knowledge is structured and applied, then refines that model as the agent works. One mechanism, Self-Directed Source Learning, has the agent flag what it still does not understand and go back to the source to fill gaps. A second, Task-Guided Source Learning, uses mistakes and patterns from actual tasks to reorganize how that knowledge is represented. Tested across five benchmarks and three different LLM backends, SourceLearn beat a Hybrid RAG baseline by up to 22.6 points and won in 13 of 15 test settings.
That gap matters because most AI agents today treat a reference document as a search index: same questions, same retrieval, no improvement over time. SourceLearn's pitch is that an agent working repeatedly with the same manual, codebase, or policy document should get measurably better at it, the way a new employee does after a few weeks on the job. For companies building agents on top of internal documentation, that distinction between access and competence could matter more than which retrieval algorithm they pick.
Of course, a 22.6 point benchmark gain is not the same as a chatbot that actually gets smarter about your tax code after the tenth question.