TGViewer
All about AI, Web 3.0, BCI All about AI, Web 3.0, BCI @alwebbci · 3.89K subscribers
Post #4425 617
Meet Q-Learning with World Models (QWM)

World models have emerged as one of the biggest directions in physical AI.

At the same time, RL fine-tuning is unlocking capabilities in frontier models beyond what pretraining can achieve on its own.

Can we get the best of both worlds?

Meet Q-Learning with World Models (QWM).

QWM isn’t tied to a specific RL algorithm it’s a general framework for test-time scaling in RL.

On top of RLPD, it also shows consistent gains.

QWM scales to frontier video models.

QWM with frontier video models also consistently improves over the base RL method
pd-perry.github.io QWM: Q-Learning with World Models QWM uses learned world models for test-time tree search in online Q-learning, improving action selection without training on imagined rollouts.
  • 🔥 1
More from @alwebbci
  1. Sep 25, 2026New From DeepSeek: a sandbox platform running 3 million AI agent environments per day. Dee…
  2. Sep 25, 2026Super interesting paper from Google and colleagues. It studies where it's possible to dist…
  3. Sep 24, 2026DeepMind Institute presented 3 new essays: 1. How can we control misbehaviour in agent swa…
  4. Sep 24, 2026Stanford introduced Matryoshka Attribution, a new attribution method which uses gradient d…
  5. Sep 23, 2026Claude has discovered a previously unknown enzyme system hidden in the DNA of bacteriophag…
  6. Sep 23, 2026Xiaomi released MiMo-V2.6 - Pro & Flash 2 omnimodal models, advancing through scaled reinf…
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →