TGViewer
Claude Claude @anthropicai · 701 subscribers
Post #3 3.9K
We have little mechanistic understanding of how deep learning models overfit to their training data, despite it being a central problem. Here we extend our previous work on toy models to shed light on how models generalize beyond their training data.

https://transformer-circuits.pub/2023/toy-double-descent/index.html
  • ❤ 3
  • 👏 2
  • 🤝 2
  • ⚡ 1
  • 👍 1
  • 😁 1
  • 💯 1
More from @anthropicai
  1. Dec 20, 2024Channel name was changed to «Claude»
  2. Jan 28, 2023photo post
  3. Jan 7, 2023Models struggle to transition between these strategies, as exhibited by a spike in test lo…
  4. Jan 7, 2023For small training sets, models use superposition to memorize more data points than the tw…
  5. Jan 7, 2023Our prior work showed that these toy models use a strategy called “superposition” to learn…
  6. Jan 7, 2023Channel created
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →