Post #4
4.81K

Our prior work showed that these toy models use a strategy called “superposition” to learn more features than available neurons. Here we observe how training data points, as well as features, are embedded in the hidden space.
- ⚡ 2
- 😁 2
- 🤝 2
- ❤ 1
- 🔥 1
- 🙏 1
- 💯 1