Hardware/ quantization · ai-hardware · semiconductors

A Chip That Lets AI Weights Mix Precision Block by Block

A new transformer chip assigns different bit-widths to blocks within a single weight matrix, cutting AI inference costs without gutting accuracy.

A chip that lets AI models mix bit-widths within a single weight matrix, not just across layers, has made it out of the lab.

The project, called ShatterQuant, pairs a quantization method with custom accelerator hardware. Instead of assigning one bit-width to an entire tensor, it splits weights into blocks and gives each block 1, 2, 4, or 8 bits based on how sensitive that block is to precision loss. The matching chip, built on TSMC's 16nm process and running at 1 GHz, reconfigures its processing elements to match each block's precision and handles softmax and other nonlinear math on-chip. In testing, it hit 1.5 TOPS of throughput, 760 GOPS per square millimeter of area, and 2.8 TOPS per watt.

On image classification with DeiT on ImageNet-1K, the approach landed within 3.3% of the accuracy of leading mixed-precision methods while using, on average, 2 fewer bits per weight - a real efficiency gain. It also held up on PixelDiT, an image-generation model, without a visible quality hit. That is the entire pitch of mixed-precision quantization: squeeze out memory and power without falling off an accuracy cliff.

Worth remembering: this is a 16nm research chip, not a product roadmap. Commercial AI accelerators have been creeping toward finer-grained precision control for years. The open question is whether fine-grained, intra-tensor quantization like this ever makes it into shipping silicon, or stays a clever idea that dies in a PDK.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →