[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-recurrent-transformer-nearly-matches-models-6x-its-size":10,"sections":40},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":35,"feedback":39,"feedback_at":22,"cost_usd":39,"total_tokens":39},6462,"recurrent-transformer-nearly-matches-models-6x-its-size","Recurrent Transformer Nearly Matches Models 6x Its Size","A shared weight recurrent transformer nearly matches models 4 to 6.4 times larger on two sequence reasoning tasks, without generating extra tokens.","A new transformer architecture skips the token-heavy chain-of-thought approach and still reasons like a much larger model.\n\nResearchers built a depth-recurrent transformer that reasons by looping a single shared-weight block rather than generating extra tokens, so each added reasoning step costs flat memory instead of growing the key-value cache the way chain-of-thought does. They tested it on three tasks with decreasing structural cues: graph reachability, nested boolean logic, and unstructured relational text. On the two sequence tasks, nested boolean logic and unstructured text, the recurrent model landed within two points of fixed-depth transformers using 4 to 6.4 times more parameters. On the graph task it extrapolated well past its training range, something fixed-depth models could barely manage, though the paper does not report a matching parameter-count comparison for that result.\n\nChain-of-thought reasoning gets expensive to serve at scale because every reasoning token bloats the key-value cache; this approach trades token generation for looping a small block, which could meaningfully cut inference memory for reasoning-heavy workloads. It also backs a broader argument that reasoning ability comes from computation depth, not raw parameter count, a distinction that matters for both training and serving budgets.\n\nThese are still narrow, synthetic compositional tasks, not open-ended reasoning, so the real test is whether the trick survives contact with messier problems like actual math or code.","[\"ai\",\"transformers\",\"reasoning\",\"research\"]","2026-09-16T04:00:00.000Z","2026-09-17T17:52:01.191Z","2026-09-17T17:52:13.088Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"The source attributes the 'within two points of 4-6.4x larger models' result to the two sequence tasks (nested boolean logic and unstructured relational text), not to 'the two structured tasks' as the headline, dek, and body claim — the graph task's extrapolation success is described separately in the source with no quantified parameter comparison, so recheck the source and correctly attribute which task grouping the parameter-efficiency claim applies to.","resolved","ai",[30,32,33,34],"transformers","reasoning","research",[36],{"name":37,"url":38},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2603.21676",0,{"sections":41},[42,46,51,56,61,65,69,74,79,83,88,93,98,103],{"name":43,"slug":30,"count":44,"latest_published_at":45},"AI",3853,"2026-09-17T08:27:09.000Z",{"name":47,"slug":48,"count":49,"latest_published_at":50},"Security","security",648,"2026-09-17T04:00:00.000Z",{"name":52,"slug":53,"count":54,"latest_published_at":55},"Policy","policy",338,"2026-09-11T04:00:00.000Z",{"name":57,"slug":58,"count":59,"latest_published_at":60},"Deals","deals",179,"2026-06-29T20:02:07.000Z",{"name":62,"slug":63,"count":64,"latest_published_at":50},"Hardware","hardware",154,{"name":66,"slug":67,"count":68,"latest_published_at":50},"Science","science",114,{"name":70,"slug":71,"count":72,"latest_published_at":73},"Consumer Tech","consumer-tech",99,"2026-09-09T17:27:33.000Z",{"name":75,"slug":76,"count":77,"latest_published_at":78},"Software","software",75,"2026-09-10T20:41:21.000Z",{"name":80,"slug":81,"count":82,"latest_published_at":50},"Dev Tools","dev-tools",73,{"name":84,"slug":85,"count":86,"latest_published_at":87},"Startups","startups",55,"2026-09-09T23:14:29.000Z",{"name":89,"slug":90,"count":91,"latest_published_at":92},"Gaming","gaming",43,"2026-09-10T12:18:06.000Z",{"name":94,"slug":95,"count":96,"latest_published_at":97},"General","general",41,"2026-09-08T01:57:23.000Z",{"name":99,"slug":100,"count":101,"latest_published_at":102},"Reviews","reviews",20,"2026-06-24T12:00:01.000Z",{"name":104,"slug":105,"count":106,"latest_published_at":107},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]