📰 Paper reading club: PPO experience report
postponed from 28/12 to 04/01 due to weather conditions!
In light of gales that fiercely spin,
We shift our date, not cowed, but keen.
With spirits high and safety first,
We'll meet anew, when winds disperse.
Primacy bias in PPO
"The Primacy Bias in RL" shows how we can improve sample efficiency and increase the expected return with weight resets. I will tell how we tried leveraging resets in PPO and analyzed primacy bias phenomena.
The plan:
- How to compare RL algorithms: "rliable" recommendations
- The Primacy Bias in RL: when resets work, and when they don't
- Resets in PPO: what we tried (spoiler: nothing worked)
- Heavy Priming experiment: overfitting on one batch at the beginning of the training. Differences between SAC and PPO, the probable cause of PPO stability.
⏱ 18:00-20:00 Thursday, 04 January
📍 F0RTHSP4CE, Khorava St, 18
💰 free/donation
🗣 @iwksx (🇬🇧/🇷🇺)
👮 @cad215
💬 event chat
Post #284
824

- ❤ 3