TGViewer
All about AI, Web 3.0, BCI All about AI, Web 3.0, BCI @alwebbci · 3.87K subscribers
Post #4477 455
Google presented Dream-RSI

The central idea is almost obvious in retrospect: that history is already a simulator.

Want to know if a different exploration strategy would have worked better? You don't need to re-run anything. You walk the same recorded tree in a different order. Every outcome is already there. Cost: zero executions.

The architecture has 3 layers:
The underlying coding agent stays completely untouched. On top of it sits a lightweight orchestration harness an executable policy that controls branching decisions, parallel exploration, and stopping rules.

Above that, a policy development agent rewrites the harness code between rounds, testing each new version against the replay simulator before any of it ever runs live.

This separation is the key design choice. The harness makes exploration programmable without touching the model underneath. And because the harness is just code, it can be evaluated, revised, and selected offline thousands of candidate versions screened against accumulated history before a single one ships.

The loop: 1. The agent explores online and builds a discovery tree 2. That tree becomes a replay simulator 3. Thousands of alternative harness policies are tested inside it no real runs, just traversals 4. The best harness is redeployed, records a new tree, expands the pool.
Dream-Rsi Dream-RSI: Recursive Self-Improvement through Evolving Worlds Accumulated discovery history becomes a replay simulator. Dream inside it to improve exploration policies at near-zero cost, then redeploy online.
  • ❤ 3
  • 🔥 3
More from @alwebbci
  1. Sep 18, 2026Researchers trained models on synthetic stories about humans only (no AIs). Team found the…
  2. Sep 17, 2026Agent harnesses seem to make a lot of difference, at least for cost, even on open source c…
  3. Sep 17, 2026S&P Global has entered an agreement to acquire smart contract security firm OpenZeppelin.…
  4. Sep 17, 2026The world's first foundation model designed for decentralized multi-robot collaboration. L…
  5. Sep 17, 2026Zhipu AI sharing how GLM-5.3 helped build and optimize the inference infrastructure servin…
  6. Sep 16, 2026Claude Cowork and chat are merging into one Claude. Ask a quick question or hand over a re…
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →