TGViewer
All about AI, Web 3.0, BCI All about AI, Web 3.0, BCI @alwebbci · 3.88K subscribers
Post #4451 683
Can Looped Transformers still help when FLOPs, parameters, and KV cache are all matched?

Bytedance introduced SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers.

The answer is yes. And the advantage grows with scale.
arXiv.org SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers Looped Transformers increase effective depth by iterating a shared block of layers, but most evaluations compare at fixed model size, conflating architectural advantage with extra FLOPs. We study...
  • 🔥 5
  • ❤ 2
More from @alwebbci
  1. Sep 25, 2026Super interesting paper from Google and colleagues. It studies where it's possible to dist…
  2. Sep 24, 2026DeepMind Institute presented 3 new essays: 1. How can we control misbehaviour in agent swa…
  3. Sep 24, 2026Stanford introduced Matryoshka Attribution, a new attribution method which uses gradient d…
  4. Sep 23, 2026Claude has discovered a previously unknown enzyme system hidden in the DNA of bacteriophag…
  5. Sep 23, 2026Xiaomi released MiMo-V2.6 - Pro & Flash 2 omnimodal models, advancing through scaled reinf…
  6. Sep 23, 2026Another day, another 20-something making a fortune from selling data to AI labs. This time…
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →