🤖 They shrunk an AI model in half and it got smarter
Multiverse Computing took OpenAI's GPT-OSS — 120 billion parameters — cut it to 60 billion, compressed memory to 4-bit, and the smaller version _outperformed_ the original on 7 of 9 benchmarks.
The trick: normally you teach the shrunken model using a halfway copy, which is already worse than the original. Their method, Quantization-Aware Healing, just points the small model at the full uncompressed original instead. That swap apparently makes all the difference.
The result uses a quarter of the memory — gap between a data center and a decent desktop. The model, Hypernova-60B, is free on Hugging Face, though the compression tool itself stays proprietary. Decrypt has more.
Post #3554
2.37K

- 👍 5
- ❤ 2