Training language models without language.
Loosely, natural data mixes contingent information (facts about our world) with universal predictive structure (composition, repetition, recursion, etc…) that are not specific to our world. Our self-play approach only supplies the latter. However, if it produces universal structure efficiently, and if universal structure is the bottleneck, predictable scaling on natural data follows.
https://arxiv.org/abs/2609.30063
https://github.com/acowsik/self_play_pretraining