AI/ state-space-models · recurrent-neural-networks · sequence-modeling · ai-research

A Tiny Recurrent Model Tackles AI's State Tracking Problem

A new complex-valued recurrent model claims near-perfect accuracy on a state-tracking task where Transformers and Mamba break down over long sequences.

A new recurrent architecture claims it can hold onto state over sequences four times longer than anything it trained on, a trick that trips up both Transformers and Mamba.

A newly posted paper describes the Complex State Propagator (CSP), a small recurrent model that tracks state using complex numbers instead of the usual real-valued vectors. Rather than running values through amplitude changes or per-step nonlinear functions, CSP converts coordinates directly into a phase angle with the atan2 function, then just keeps rotating that phase forward through the network. The authors tested it on Mod-3 tracking, a benchmark where a model has to correctly track a running value mod 3 as it reads through a sequence. A 3-layer version trained on sequences of length 16 reportedly held near 100% accuracy and F1 score when tested on sequences four times longer, at length 64.

Deterministic state tracking - remembering a precise running count or parity as a sequence gets longer - is a known weak spot for Transformers and Mamba, both of which tend to drift once sequences run past anything seen in training. It is also a decent proxy for whether a model is doing real algorithmic reasoning instead of pattern-matching. If a model this small holds up on out-of-distribution length the way the paper claims, that is worth watching, especially since it did not need a large parameter count to get there.

That said, the paper leans hard on phrases like 'architectural marvel' and 'absolute mathematical purity' for a result tested on one synthetic counting task - real language and reasoning workloads are a messier test it does not attempt.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →