AI/ ai · model-architecture · attention-mechanism · research

AMOR Model Saves Compute by Skipping Attention When Sure

AMOR activates attention only when a language model's uncertainty spikes, cutting compute while topping fixed-schedule hybrids on reasoning benchmarks.

A new hybrid AI architecture skips expensive attention computation unless the model is actually confused.

Researchers behind AMOR (Adaptive Metacognitive Output Router) built a recurrent backbone, using Mamba2 or Gated DeltaNet, and bolted on attention blocks that only switch on when the model's output entropy crosses a dynamic threshold set from a running batch median and standard deviation. No learned routing parameters are involved; it is a straightforward binary gate. Trained from scratch on FineWeb-Edu, the approach invoked attention on roughly 40% of positions yet still posted the best eight-task common-sense reasoning average at each model scale, beating pure recurrent, pure attention, and fixed-schedule hybrid baselines. It also improved retrieval over recurrent-only models, held its own against fixed-schedule hybrids, and kept the long-context stability of its recurrent backbone under distribution shift, a spot where Transformers and other hybrids reportedly degrade.

The result challenges the assumption that hybrid models need attention on a fixed schedule or a learned router to perform well. Timing attention to when the model's own uncertainty spikes, rather than deciding how often to apply it in advance, seems to matter more for both accuracy and efficiency. For teams building long-context or resource-constrained models, that is a far cheaper lever to pull than redesigning the whole architecture.

It is a clean result, but it comes from smaller-scale pretraining runs on curated data and a fixed benchmark suite, so whether the gains hold at production-scale training is still an open question.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →