“Intel researchers just found a way to push ternary LLMs below the conventional 1.58-bit barrier, without changing a single weight.”
AI is slowly escaping the datacenter.
& Compression breakthroughs like this are part of how it gets into everything else.👀
Intel researchers just found a way to push ternary LLMs below the conventional 1.58-bit barrier, without changing a single weight.
Their new BITCOS method exploits the unusually high number of zero weights in ternary models.
Across 29 ternary LLM checkpoints, zero weights reached up to 51.48%.
That allowed BITCOS to reach just 1.485 bits per weight, while preserving the exact ternary weights.
In testing, it delivered up to 18% higher CPU decode throughput and up to 27% higher GPU decode throughput.
Smaller models. Less memory movement. Faster inference.
This could become increasingly important for running powerful AI on PCs, smartphones, robots and edge devices.