[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-mosar-teaches-language-models-where-attention-actually-matters":10,"sections":40},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":35,"feedback":39,"feedback_at":22,"cost_usd":39,"total_tokens":39},8059,"mosar-teaches-language-models-where-attention-actually-matters","MoSAR teaches language models where attention actually matters","A new attention mechanism learns when language models need broad context, with early signs it could also cut inference costs down the line.","MoSAR is a new attention mechanism that lets language models decide, on their own, which words deserve a global view and which just need a glance at their neighbors.\n\nResearchers behind the paper introduce query and key routers, placed right after positional encoding, that pick from short, medium, and global attention regimes for each input. Instead of hard-coding a sparsity pattern the way many efficient-attention schemes do, MoSAR produces a continuous, distance-dependent attention field that's learned during training. In controlled pretraining runs with matched 500-million-parameter models, MoSAR learned a narrower attention geometry without hurting language-modeling quality, and it beat dense RoPE attention on perplexity at the training context length. When tested on longer contexts than it trained on, it beat every baseline evaluated, including ALiBi, a well-regarded approach for context extrapolation.\n\nThe quadratic cost of standard self-attention is the reason long-context models get expensive fast, and most fixes so far have guessed in advance where attention can be sparse. MoSAR's pitch is that the model should figure that out itself, since which words actually relate to which other words depends on the sentence, not a fixed rule. The paper also found the learned attention pattern holds up when simplified to a discrete top-1 choice after training, hinting, but not proving, that it could run cheaper at inference too.\n\nThis is a 500M-parameter proof of concept, several orders of magnitude smaller than the models actually straining under long-context costs today, and the paper measures perplexity, not wall-clock speed or memory savings; efficient-attention ideas have a long history of looking great in academic experiments and then quietly not shipping in production models. MoSAR still has that gauntlet to run.","[\"ai\",\"attention-mechanisms\",\"language-models\",\"research\"]","2026-09-28T04:00:00.000Z","2026-09-28T07:46:55.700Z","2026-09-28T07:47:02.395Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"The dek claims MoSAR works 'without the usual quadratic cost' as settled fact, but the body itself only hedges that efficiency gains 'could show up' at inference time via discretization — the source paper demonstrates perplexity improvements, not a proven reduction in quadratic attention cost, so reword the dek to match the more tentative claim the body and source actually support.","resolved","ai",[30,32,33,34],"attention-mechanisms","language-models","research",[36],{"name":37,"url":38},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.31261",0,{"sections":41},[42,45,49,54,59,64,68,73,78,83,88,93,97,102],{"name":43,"slug":30,"count":44,"latest_published_at":18},"AI",4750,{"name":46,"slug":47,"count":48,"latest_published_at":18},"Security","security",759,{"name":50,"slug":51,"count":52,"latest_published_at":53},"Policy","policy",399,"2026-09-27T18:39:02.000Z",{"name":55,"slug":56,"count":57,"latest_published_at":58},"Deals","deals",261,"2026-09-27T15:30:35.000Z",{"name":60,"slug":61,"count":62,"latest_published_at":63},"Hardware","hardware",188,"2026-09-27T20:46:36.000Z",{"name":65,"slug":66,"count":67,"latest_published_at":18},"Science","science",151,{"name":69,"slug":70,"count":71,"latest_published_at":72},"Consumer Tech","consumer-tech",135,"2026-09-26T14:30:00.000Z",{"name":74,"slug":75,"count":76,"latest_published_at":77},"Software","software",91,"2026-09-25T20:55:00.000Z",{"name":79,"slug":80,"count":81,"latest_published_at":82},"Dev Tools","dev-tools",84,"2026-09-26T04:20:58.000Z",{"name":84,"slug":85,"count":86,"latest_published_at":87},"Startups","startups",76,"2026-09-25T18:33:59.000Z",{"name":89,"slug":90,"count":91,"latest_published_at":92},"Gaming","gaming",48,"2026-09-25T18:35:21.000Z",{"name":94,"slug":95,"count":91,"latest_published_at":96},"General","general","2026-09-26T17:02:42.000Z",{"name":98,"slug":99,"count":100,"latest_published_at":101},"Reviews","reviews",30,"2026-09-24T20:07:31.000Z",{"name":103,"slug":104,"count":105,"latest_published_at":106},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]