TGViewer
Data Science Data Science @datascienceiot · 42.6K subscribers
Post #3780 567
Shockingly Simple Self-retrospection Improves Agentic Models Without RL

"Can a language-model agent improve its future actions by training only on explanations of its own experience? We investigate this question by studying Retrospection-Only Fine-Tuning (ROFT), a minimal online procedure designed to isolate the effect of explanation-only training on subsequent behavior. The agent attempts a task, observes available feedback, generates a retrospective explanation, and is fine-tuned with a next-token prediction loss on the explanation tokens alone."

📓 Read

@datascienceiot
More from @datascienceiot
  1. Sep 30, 20265 октября начнется 19-й поток программы Data Engineer от Newprolab Программа для junior- и…
  2. Sep 28, 2026"PrimeScientist: Strategic Allocation of Research Effort in Autonomous Research" 📓 Read @…
  3. Sep 28, 2026ML-проект не заканчивается на обучении модели В Data Science есть довольно знакомый момент…
  4. Sep 25, 2026New Stanford+Together AI paper shows teams that learn to check each other's work solve pro…
  5. Sep 25, 2026Avito DS Meetup: RL, AI-агенты и BidCorrection 🚀 🔸 Булат Шарипов расскажет, как в Авито…
  6. Sep 24, 2026Agent Memory Is a Surface for Endogenous Authorization Laundering 📘 Paper @datascienceiot
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →