Samsung's research arm just published a way to compress large language model weights to less than one bit each.
The project, called LittleBit, is live on GitHub under Samsung Labs. It uses a technique the team calls latent factorization to push weights below the one-bit-per-parameter mark that most aggressive quantization schemes treat as a practical floor. The repository surfaced on Hacker News on October 8, 2026, drawing 37 points but only four comments, a muted reaction for a claim this bold. The listing we have doesn't include a paper, benchmark numbers, or which model sizes were tested.
Squeezing a model's weights further is the main lever for running LLMs on a phone or laptop instead of a server rack. Most previous sub-2-bit efforts, including Microsoft's BitNet line, have settled around 1.58 bits per weight as the point where accuracy still holds up; going under a full bit would be a real step if LittleBit's accuracy survives the squeeze.
Compression numbers are cheap to claim and expensive to verify. Until someone runs LittleBit against a standard benchmark and model, file this under promising, not proven.