A new recurrent architecture claims it can hold onto state over sequences four times longer than anything it trained on, a trick that trips up both Transformers and Mamba.
A newly posted paper describes the Complex State Propagator (CSP), a small recurrent model that tracks state using complex numbers instead of the usual real-valued vectors. Rather than running values through amplitude changes or per-step nonlinear functions, CSP converts coordinates directly into a phase angle with the atan2 function, then just keeps rotating that phase forward through the network. The authors tested it on Mod-3 tracking, a benchmark where a model has to correctly track a running value mod 3 as it reads through a sequence. A 3-layer version trained on sequences of length 16 reportedly held near 100% accuracy and F1 score when tested on sequences four times longer, at length 64.
Deterministic state tracking - remembering a precise running count or parity as a sequence gets longer - is a known weak spot for Transformers and Mamba, both of which tend to drift once sequences run past anything seen in training. It is also a decent proxy for whether a model is doing real algorithmic reasoning instead of pattern-matching. If a model this small holds up on out-of-distribution length the way the paper claims, that is worth watching, especially since it did not need a large parameter count to get there.
That said, the paper leans hard on phrases like 'architectural marvel' and 'absolute mathematical purity' for a result tested on one synthetic counting task - real language and reasoning workloads are a messier test it does not attempt.