TGViewer
Claude Claude @anthropicai · 701 subscribers
Post #7 8.25K
Models struggle to transition between these strategies, as exhibited by a spike in test loss. This spike moves to larger datasets as one increases model capacity. This is a clear signature of double-descent, a phenomenon that is now well-known in the ML literature.

We hope these results are a step towards a mechanistic theory of memorization. There are many open questions, such as understanding the loss spike, or what happens when only a subset of the data is repeated.

Thanks to Adam Jermyn for his comments reproducing and extending these results!

https://transformer-circuits.pub/2023/toy-double-descent/index.html#comment-jermyn-1
  • ❤ 2
  • 👍 2
  • 👏 2
  • 🙏 2
  • 😁 1
  • 😍 1
  • 🤝 1
More from @anthropicai
  1. Dec 20, 2024Channel name was changed to «Claude»
  2. Jan 28, 2023photo post
  3. Jan 7, 2023For small training sets, models use superposition to memorize more data points than the tw…
  4. Jan 7, 2023Our prior work showed that these toy models use a strategy called “superposition” to learn…
  5. Jan 7, 2023We have little mechanistic understanding of how deep learning models overfit to their trai…
  6. Jan 7, 2023Channel created
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →