A new AI research harness keeps a running lab notebook so it never has to choose between forgetting useful experiments and drowning in them.
Researchers built LabBook, a single-agent system for AI-driven discovery tasks. Typical evolutionary approaches generate new candidate programs from a small set of selected ancestors, which keeps context short but can miss useful evidence sitting elsewhere in the run. LabBook instead has one agent keep a notebook that both guides retrieval from the full experimental log and informs the next program it writes, updating both together at every step. Across 49 Frontier-CS problems, it beat evolutionary baselines on the quality-cost trade-off with two different model backbones, and held its own on nine more math, systems, and heuristic-design tasks.
This is really a memory-management fix, not a new kind of intelligence. Evolutionary LLM search methods control cost by only showing the model a handful of ancestor programs, which works until the useful lesson lives in an experiment that was never sampled. LabBook's bet is that letting the agent write its own running summary is a cheaper way to avoid that blind spot without paying for the full transcript every time.
The gains reported here are efficiency gains, not new capability: the 49 benchmark problems were already solvable, just more expensively. Whether a self-written notebook holds up on messier, real-world discovery problems is the open question, and the code is not out yet to check.