An RL method that solves the key problem of unstable training in LLMs and MoE architectures and offers a more reasonable and gentle approach to controlling the learning process.
🟢 How it Solves the problem?
Reinforcement Learning, RL - is the ingredient that turns a simple large language model into a reasoning assistant. It's RL that teaches AI to solve olympiad math problems, write clean code, and understand the relationship between text and images.
But RL has a downside: catastrophic training instability, especially for gigantic models.
The main technical puzzle is controlling the importance coefficients at the level of each token. In MoE architectures, where different parts of the model are activated for different tasks, these coefficients can "jump" uncontrollably.
Excessively large fluctuations in the coefficients turn clear training signals into interference, destabilizing the entire system.
Until now, the standard tools have been GRPO and GSPO, which used the principle of hard clipping. If the coefficient exceeded the specified limits, the gradient was simply set to zero.
FIRST MINUS: Information loss. Valuable but outlier data were ruthlessly discarded.
SECOND MINUS: Impossible balance. If you make the limits too narrow, you stifle learning. If you make them too wide, parasitic noise creeps in. For capricious MoE architectures, this dilemma is particularly relevant.
SAPO proposes to abandon hard clipping in favor of intelligent smoothing.
Instead of abruptly setting the gradient to zero, SAPO uses a smooth, adaptive function (controlled by temperature) that gently reduces the influence of problematic gradients without completely nullifying them. This creates continuous confidence regions within which the model can learn more flexibly and safely.
Like GSPO, but smarter. If only one token in a long answer was wrong, GSPO punished the entire sequence. SAPO selectively suppresses only the "culprit", preserving useful signals from the rest of the words. This dramatically improves the efficiency of training data sets.
Like GRPO, but smoother. Instead of abruptly disabling the gradient for a bad token, SAPO applies a gradual attenuation. This prevents sudden jumps in learning and ensures a smooth and stable adjustment of the model's policy.
The icing on the cake of the method is an asymmetric temperature design. SAPO processes "good" and "bad" updates differently. For tokens with a negative contribution, a higher temperature is used, forcing their influence to attenuate faster and stronger.
This simple rule reliably suppresses the most dangerous fluctuations, which in practice leads to unprecedented stability of the RL learning process.
CONFIRMED BY TESTS
When training Qwen3-30B-A3B-Base, SAPO not only showed a more stable learning curve, but also achieved higher results on complex mathematical benchmarks AIME25, HMMT25. And he did this without the labor-intensive routing reproduction that competitors needed to work with MoE.
The success was repeated in a large-scale experiment with multimodal Qwen3-VL-30B-A3B, where SAPO consistently outperformed its analogues in mixed tasks on coding, logic, and mathematics.
#AI #ML #LLM #MoE #SAPO #Qwen
••••••••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
