We have little mechanistic understanding of how deep learning models overfit to their training data, despite it being a central problem. Here we extend our previous work on toy models to shed light on how models generalize beyond their training data.
https://transformer-circuits.pub/2023/toy-double-descent/index.html
Post #3
3.9K

- ❤ 3
- 👏 2
- 🤝 2
- ⚡ 1
- 👍 1
- 😁 1
- 💯 1