TGViewer
Claude Claude @anthropicai · 711 subscribers
Post #6 5.93K
For small training sets, models use superposition to memorize more data points than the two available neurons. For large training sets, models learn features in superposition, as observed in our previous work, allowing the model to generalize.

https://transformer-circuits.pub/2022/toy_model/index.html
  • 😁 3
  • 🙏 3
  • ❤ 2
  • 👍 2
  • 🔥 1
  • 👏 1
  • 😍 1
  • 💯 1
More from @anthropicai
  1. Dec 20, 2024Channel name was changed to «Claude»
  2. Jan 28, 2023photo post
  3. Jan 7, 2023Models struggle to transition between these strategies, as exhibited by a spike in test lo…
  4. Jan 7, 2023Our prior work showed that these toy models use a strategy called “superposition” to learn…
  5. Jan 7, 2023We have little mechanistic understanding of how deep learning models overfit to their trai…
  6. Jan 7, 2023Channel created
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →