A new paper proposes training neural networks whose internal layers are organized into geometric shapes called polytopes, so each layer's structure can be read directly instead of reverse-engineered after the fact.
The method, called Polytopal Neural Networks, bakes that shape into the network during training rather than bolting on an explanation after the fact. Each input is described by how closely it aligns with layer-specific facets of its polytope. The researchers scaled this with learned corpus representations and an inference procedure that maps activations onto a simplex, and found the same setup doubles as a new way to run vector-quantized training.
Most interpretability work either pokes at a finished model from the outside or relies on a fixed dataset to define what a concept means. This approach makes the structure part of how the network actually computes, and the paper reports that comes with little to no accuracy cost while producing cleaner compressed representations than standard vector quantization.
Interpretability methods have a habit of looking tidy in a paper and then getting messy on real production models, so the real test is whether polytopal constraints hold up outside the benchmarks reported here.