i wonder when we'll get to training over binary ops (and if this is fundamentally possible) bc 1-bit as shown in the article is more like dictionary training, the matrices still contain real numbers
https://arxiv.org/abs/2502.05003
Post #862
238