📰 Paper reading club: Scalable Supervision Potpourri
We meet biweekly to discuss advances in computer science, ground them down to applications, and see where to the wind is blowing. Read the papers at your own time and come discuss them with us. The talks will be given in English, discussion in En/Ru.
12:00 Scalable Supervision (@add512)
Can we scale human feedback for complex AI tasks? Can weak-and-safe AIs oversee strong-and-unsafe ones? We'll start with an introduction and cover (some of) the following papers. Please read at least one.
- Supervising strong learners by amplifying weak experts
- AI safety via debate
- Weak-to-Strong Generalization: Eliciting Strong Capabilities with Weak Supervision
- Preventing Language Models from Hiding Their Reasoning
- AI Control: Improving Safety Despite Intentional Subversion
⏱ 12:00-14:30 Saturday, 13 Apr
📍 F0RTHSP4CE, Khorava St, 18
🗓 please add 🫡 below if you're coming
💰 free/donation/boost
👅 en+ru
👮 @add512
💬 event chat and q&a
🗣 propose talk
🎥 recordings
Post #354
1.73K

- 🫡 5
- ❤ 2
- 👍 1
- 👀 1