A new academic framework squeezes high-resolution 3D shape generation into far fewer computational tokens without losing structural accuracy.
Researchers published SILSA (Sliding-Window Slice Latents), detailed in an arXiv paper posted October 2, 2026. Instead of the voxel-token pipelines most high-resolution 3D generators rely on, SILSA represents shapes as overlapping slices along three axes, with each slice summarizing a local depth window. A Slice VAE encodes surface samples into these slice latents, a sparse volumetric decoder reconstructs them, and a Volumetric Anchor Lattice coordinates the directional slice streams in a shared 3D workspace. The team also added topology-specific supervision - matching persistence diagrams and Betti-number transitions between neighboring slices - so thin or highly connected shapes do not fall apart during generation.
The reported numbers: 8.7% better PSNR, a 5.96 point gain in coverage, and 9.2% lower Betti error than the strongest baseline, achieved using 70% fewer tokens than the next most compact method and over 98% fewer than sparse or hierarchical tokenizers. That cut training memory by 40.4% and inference time by 58.5%. For an industry trying to make 3D generation practical for game assets, CAD, and robotics simulation, the real bottleneck has long been compute cost paired with geometry that looks fine until you check whether a loop or thin strut actually stays connected.
Most 3D generators chase realism. This one is making the more boring but arguably more useful case - that structure should survive the generation process at all.