[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-new-chip-fabric-design-cuts-ai-memory-traffic-by-up-to-58":10,"sections":35},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":24,"tags":25,"sources":30,"feedback":34,"feedback_at":22,"cost_usd":34,"total_tokens":34},6830,"new-chip-fabric-design-cuts-ai-memory-traffic-by-up-to-58","New Chip Fabric Design Cuts AI Memory Traffic by Up to 58%","A new network-on-chip fabric called MeshKV reroutes AI chip memory traffic to ease a bottleneck that limits how fast models generate long responses.","A new interconnect design promises to unclog one of the biggest bottlenecks in AI inference: moving cached data around inside the chip.\n\nResearchers behind a system called MeshKV built a network-on-chip fabric that treats key-value cache data (the running memory a transformer model uses to generate each new token) as packetized traffic routed over a lightweight on-chip network, rather than funneling everything through a handful of centralized memory paths. The design combines three techniques: affine striping that spreads cache data across memory banks to avoid hotspots, multicast routing with duplicate-suppression to cut redundant transfers, and pipelining that overlaps data prefetch with computation. The team tested it on an 8x8 FPGA running LLaMA-2-7B and Mistral-7B at context lengths of 8K to 32K tokens. Interconnect traffic dropped by up to 58%, KV cache bandwidth utilization improved 2.1x, and throughput for multiple simultaneous streams rose as much as 1.9x.\n\nThe bottleneck this targets is real and growing. As context windows stretch into tens of thousands of tokens, the KV cache, not raw compute, increasingly determines how fast and how many requests an accelerator can serve at once. Most fixes to date have focused on compressing the cache or shuffling it between DRAM tiers; MeshKV instead rethinks the wiring that moves it, treating a decoding accelerator more like a small data-center network than a monolithic chip.\n\nIt is still an FPGA prototype, not silicon shipping in anyone's data center, so the gains here describe a research testbed, not a production inference stack. The real test is whether this kind of on-chip network survives the jump to full ASIC accelerators at the scale hyperscalers actually run.","[\"ai chips\",\"interconnect design\",\"kv cache\",\"llm inference\"]","2026-09-18T04:00:00.000Z","2026-09-18T18:35:37.218Z","2026-09-18T18:35:49.160Z","published",null,[],"hardware",[26,27,28,29],"ai chips","interconnect design","kv cache","llm inference",[31],{"name":32,"url":33},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.19207",0,{"sections":36},[37,41,45,50,55,58,62,67,71,76,81,86,91,96],{"name":38,"slug":39,"count":40,"latest_published_at":18},"AI","ai",4031,{"name":42,"slug":43,"count":44,"latest_published_at":18},"Security","security",654,{"name":46,"slug":47,"count":48,"latest_published_at":49},"Policy","policy",338,"2026-09-11T04:00:00.000Z",{"name":51,"slug":52,"count":53,"latest_published_at":54},"Deals","deals",179,"2026-06-29T20:02:07.000Z",{"name":56,"slug":24,"count":57,"latest_published_at":18},"Hardware",155,{"name":59,"slug":60,"count":61,"latest_published_at":18},"Science","science",121,{"name":63,"slug":64,"count":65,"latest_published_at":66},"Consumer Tech","consumer-tech",99,"2026-09-09T17:27:33.000Z",{"name":68,"slug":69,"count":70,"latest_published_at":18},"Dev Tools","dev-tools",78,{"name":72,"slug":73,"count":74,"latest_published_at":75},"Software","software",75,"2026-09-10T20:41:21.000Z",{"name":77,"slug":78,"count":79,"latest_published_at":80},"Startups","startups",55,"2026-09-09T23:14:29.000Z",{"name":82,"slug":83,"count":84,"latest_published_at":85},"Gaming","gaming",43,"2026-09-10T12:18:06.000Z",{"name":87,"slug":88,"count":89,"latest_published_at":90},"General","general",41,"2026-09-08T01:57:23.000Z",{"name":92,"slug":93,"count":94,"latest_published_at":95},"Reviews","reviews",20,"2026-06-24T12:00:01.000Z",{"name":97,"slug":98,"count":99,"latest_published_at":100},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]