It really does seem pre-training as we know it is beginning to die. Of course, Ilya spoke on this. The pre-training corpus has been exhausted. Now the primary model gains come from RL-able tasks. Of all the people at the labs I speak with, almost no one mentions pre-training anymore; just post-training/RL and continuous learning (test-time-training).
This is why the models have become so much better (on evals at least) with math/coding while seemingly having stagnated on other more general, less RL-able domains.
This is why all the labs are spending so much money on different RL envs; so they can smooth out the model capability 'spikiness'.
🄳🄾🄾🄼🄿🤖🅂🅃🄸🄽🄶
Post #267010
617

- 💯 1