Researchers have built an AI memory system that compresses a patient's entire visit history into a compact, constantly updated summary instead of rereading every record from scratch.
The system, called ReLMem, pairs a frozen large language model with lightweight adapters that update a fixed-size memory each time a new clinical visit comes in. Rather than reprocessing a patient's full history for every prediction, the model folds new visit data into the existing memory using a training method that aligns compressed and full-history attention outputs on the same queries. On a medication-prediction task built from electronic health records, ReLMem approached the F1 scores of a baseline that kept the complete history, while cutting stored data by 97.1 percent. Against the strongest existing compressed-memory approach under the same storage budget, it improved macro-F1 by 4.66 points and micro-F1 by 4.75 points.
The pitch here is cost, not magic: feeding a model someone's full, years-long medical record on every query gets expensive fast, and that expense compounds as patient histories grow. A memory that updates incrementally, without rereading everything each time, is the kind of infrastructure change that makes longitudinal EHR modeling practical to deploy rather than just publish.
The gap between 'approaches' and 'matches' full-history accuracy is the real number to watch. Compression research has a habit of quietly narrowing that gap in follow-up papers, and this reads like version one for patient memory built this way.