A new technique lets large language models borrow computation from earlier conversation turns to solve new, unrelated problems more accurately.
Researchers studied whether retained conversation history helps or hurts a model solving a later, independent problem. They found that old context can raise or lower accuracy within the same domain, even when the new problem had nothing to do with the old one. Using controlled replay experiments, they isolated how each problem-history pairing changes a model's internal state, and found those changes preserve similar relationships across different histories. Building on that, they introduced STAIR, which stores keys and values from earlier answers in a fixed bank and trains a small routing mechanism to decide when a new query should read from that bank. The base model's weights never change; only 12,288 extra parameters are trained. Tested on three Qwen models across four benchmarks, STAIR raised later-turn accuracy by up to 11.67 percentage points over an unmodified model with the same history.
Most chatbot products already keep a running context window, but that history is just text the model re-reads, not stored computation it can selectively reuse. STAIR points at a cheaper middle path between full fine-tuning and starting fresh each turn: a tiny, swappable add-on that recycles past computation instead of recomputing or ignoring it. For products juggling long sessions, like coding assistants or customer support bots, that could mean fewer wasted tokens and fewer cases where leftover context quietly degrades a new answer.
The gains come only from Qwen models and a mere 12,288 trained parameters, so whether this generalizes to the larger, more guarded models behind most consumer chatbots remains unproven.