Modern AI is still directly dependent on human labeling and human data in general. And there are a lot of problems with this: it's expensive, time-consuming, "data runs out", etc.
🟢 What META done?
They are also convinced that this is essentially a hard ceiling on the path to AGI: if you only train agents on human traces, then the learning boils down to refining human experience. So can we be 100% sure that such systems can learn something outside the distribution and become smarter than us? This is especially true for areas such as coding, which will be discussed further.
The researchers proposed Self-Play SWE-RL - a way to train agents so that they can self-improve on their own data.
Self-Play SWE-RL consists of two entities: Bug-injector and Bug-solver. The system receives a repository of code, and the Bug-injector studies it, breaks the code, and weakens the tests so that the bug can hide.
The task of Bug-solver is obvious: to fix the code, without issue-text, without hints, without ready-made test runners. And if he breaks something in the process, this case also becomes part of the dataset and expands the sample.
It's important to understand that these are not just synthetic bugs. Here, the same policy breaks and fixes the code (that is, these are just different roles of one agent). In this sense, the approach somewhat resembles GAN: the solver learns at the expense of the injector becoming smarter, and vice versa.
The results are as follows:
- Code World Model (CWM) on 32B, which has already passed the sft stage and was trained in this way, achieved +10.4% on SWE-bench Verified and +7.8% on SWE-bench Pro
- Compared to conventional RL, this approach gives +2.4% on SWE-bench Verified and +3.6% on SWE-bench Pro
Not a breakthrough, but few pipelines today give such significant increases, so it's quite interesting (but code wasn't provided).
••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
