TGViewer
RIML Lab RIML Lab @rimllab ยท 3.25K subscribers
Post #255 3.62K
๐Ÿค– RL Journal Club

โœ… This Week's Presentation:

๐Ÿ”น Title: Test Time Exploration to Achieve Generalization in Zero-Shot RL
๐Ÿ”ธ Presenter: Alireza Farajtabrizi

๐ŸŒ€ Abstract:
This paper studies zero-shot generalization in reinforcement learning, where an agent is trained on a set of tasks but must perform well on unseen test environments. The authors argue that standard reward-maximizing RL agents can overfit to training tasks, especially in environments where simple invariance-based methods fail. Their key insight is that exploration behavior is harder to memorize than reward-seeking behavior and can therefore generalize better.

To build on this idea, the paper introduces Explore to Generalize (ExpGen), an algorithm that combines a maximum-entropy exploration policy with an ensemble of reward-seeking agents. At test time, when the ensemble agrees on an action, the agent exploits that decision; when the ensemble is uncertain, the agent switches to the exploration policy to reach new parts of the state space. Experiments on ProcGen show strong improvements on challenging tasks such as Maze and Heist, setting new state-of-the-art results in several zero-shot RL settings.

๐Ÿ“„ Paper: Explore to Generalize in Zero-Shot RL (NeurIPS 2023)

Session Details:

* ๐Ÿ“… Date: Tuesday ุณู‡โ€Œุดู†ุจู‡
* ๐Ÿ•’ Time: 15:30 - 16:30
* ๐ŸŒ Location: Online at vc.sharif.edu/ch/rohban (http://vc.sharif.edu/ch/rohban)

We look forward to your participation! โœŒ๏ธ
More from @rimllab
  1. Sep 27, 2026๐Ÿ” ML Security Journal Club โœ… This Week's Presentation: ๐Ÿ”น Title: Exploiting LLM Quantizatโ€ฆ
  2. Sep 20, 2026we are looking for Teaching Assistants to join the Multi-Agent Reinforcement Learning (MARโ€ฆ
  3. Sep 20, 2026๐Ÿ” ML Security Journal Club โœ… This Week's Presentation: ๐Ÿ”น Title: Robust Machine Unlearninโ€ฆ
  4. Sep 15, 2026๐Ÿ”˜ Open Research Position: Machine Unlearning ร— Model Quantization We are looking for motiโ€ฆ
  5. Sep 13, 2026๐Ÿ” ML Security Journal Club โœ… This Week's Presentation: ๐Ÿ”น Title: Catastrophic Failure ofโ€ฆ
  6. Aug 30, 2026๐Ÿ“ข Join the IABI TA Team! ๐Ÿฉป The Intelligent Analysis of Biomedical Images (IABI) course iโ€ฆ
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook โ†’Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 โ†’