TGViewer
Parallel Experiments Parallel Experiments @linghaoch · 1.75K subscribers
Post #1009 1.32K
https://arxiv.org/abs/2503.02113

The core idea:

Deep learning does not work because neural nets somehow escape generalization theory. It works because very flexible models can still generalize when they have soft inductive biases — preferences for simple, compressible, structured solutions.

Key points:

- 🧠 Overparameterization is not automatically a problem.
Having more parameters than data points does not necessarily mean the learned function is complex. Parameter count is a bad proxy for the complexity of the actual solution.

- 📈 Benign overfitting is not unique to neural nets.
Models can perfectly fit training data, even noisy data, while still generalizing on structured data. Similar behavior appears in linear models, Gaussian processes, high-degree polynomials, and other classical model classes.

- 🔁 Double descent is not just a modern deep learning anomaly.
The pattern where test error falls, rises, then falls again as model size increases also appears outside neural networks. It can be understood through effective dimensionality, compression, and the geometry of learned solutions.

- 📦 Compression is central.
A huge model can generalize if the solution it finds is simple or compressible. The rough intuition is:
expected error ≈ training error + complexity/compressibility penalty.

- 📚 Some older theories already help explain this.
PAC-Bayes and countable hypothesis bounds are more useful here than raw VC dimension, Rademacher complexity, or parameter counting, because they focus on which solutions are likely/simple rather than just how large the hypothesis space is.

- 🎯 The paper’s recommended lens:
Don’t only restrict what a model can represent. Instead, allow a very rich hypothesis space, but bias the learner toward simpler solutions that fit the data.

- ✨ What is still distinctive about deep learning?
Not overparameterization or double descent by themselves, but things like representation learning, in-context learning, broad cross-domain usefulness, and mode connectivity in loss landscapes.

My takeaway:

Deep learning’s famous generalization puzzles may not require rewriting the textbooks from scratch. They may require reading the right parts of the textbooks more carefully — especially the parts about priors, compression, and soft preferences over solutions.
arXiv.org Deep Learning is Not So Mysterious or Different Deep neural networks are often seen as different from other model classes by defying conventional notions of generalization. Popular examples of anomalous generalization behaviour include benign...
More from @linghaoch
  1. Oct 1, 2026https://56k.rip/
  2. Sep 11, 2026https://sunilpai.dev/posts/the-task-isnt-the-job/
  3. Jul 25, 2026https://alexzhang13.github.io/blog/2026/harness/ > A good harness is a harness that reduce…
  4. Jul 6, 2026https://linghao.io/posts/taxonomy-differences-matter 以前觉得 taxonomy 只是无聊的分类学,开始做 LLM qualit…
  5. Jun 28, 2026关于层出不穷的各式 AI memory system 的一些思考:我们应该把更多的精力放在设计更好的 eval 上,从而让最强的 memory system 进化出来 https:…
  6. Apr 18, 2026一月底最后一个周六有了一个灵感,想做个解放双手,优化了 AirPods 录音,边散步边和自己对话的 App。打开 Cursor coding 了一天,第二天就出门去 SoHo 散步…
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →