Microsoft researchers have developed a new large language model (LLM) using a ternary (-1, 0, 1) weight system, significantly reducing memory footprint and computational needs. This "1.58-bit" model, trained natively at scale with 4 trillion tokens, achieves performance comparable to full-precision models, running efficiently on a desktop CPU. Unlike previous quantization methods, this open-source model avoids significant performance degradation, marking a key advancement in efficient LLM design.
Prepared by Jonathan Pierce and reviewed by editorial team.
Comments