🤖 RL Journal Club
✅ This Week's Presentation:
🔹 Title: Is Reinforcement Learning Really Harder Than Bandits?
🔸 Presenter: Arshia Gharooni
🌀 Abstract:
Episodic reinforcement learning, despite having longer planning horizons, presents little additional sample complexity difficulty compared to contextual bandits, with the proposed Monotonic Value Propagation (MVP) algorithm achieving near-optimal regret bounds. The MVP algorithm utilizes a simplified, variance-aware bonus to achieve superior performance, offering an exponential improvement in horizon dependency and sample efficiency over previous state-of-the-art methods.
📄 Paper: Is Reinforcement Learning More Difficult Than Bandits? A Near-optimal Algorithm Escaping the Curse of Horizon
Session Details:
- 📅 Date: Wednesday (چهارشنبه)
- 🕒 Time: 17:00 - 18:00
- 🌐 Location: Online at http://vc.sharif.edu/ch/rohban
We look forward to your participation! ✌️
Post #256
5.04K