Apple published & then quickly removed ⚠️ brought you a leaked stuff 😎
RLAX js a scalable reinforcement learning framework for LLMs on TPUs.
🚫 What's inside RLAX?
- Parameter server architecture
- The central trainer updates weights
- Huge inference fleets pull up weights and generate rollouts
- Optimized for preemption and massive parallelism
- Special techniques for data curation and alignment
The results are impressive:
- +12.8% to pass@8 on QwQ-32B
- In just 12 hours and 48 minutes
- 1024 TPU v5p were used
Why this is important?
- Apple is clearly experimenting with RL on a very large scale
- The TPU-oriented architecture indicates a focus on efficiency, not just the model
- The gain is achieved not through "model magic", but through engineering the training system
- This is another signal that RL for LLMs is entering the phase of industrial pipelines
What you think, Why Apple removed?
🤖 Data Science, ML & Big Data with @DataXplore
