TGViewer
Data eXplore : Data Science, ML, Big Data, LLMs and AI Security Data eXplore : Data Science, ML, Big Data, LLMs and AI Security @dataxplore · 583 subscribers
Post #1904 196
Soft Adaptive Policy Optimization aka SAPO

An RL method that solves the key problem of unstable training in LLMs and MoE architectures and offers a more reasonable and gentle approach to controlling the learning process.

🟢 How it Solves the problem?

Reinforcement Learning, RL - is the ingredient that turns a simple large language model into a reasoning assistant. It's RL that teaches AI to solve olympiad math problems, write clean code, and understand the relationship between text and images.

But RL has a downside: catastrophic training instability, especially for gigantic models.

The main technical puzzle is controlling the importance coefficients at the level of each token. In MoE architectures, where different parts of the model are activated for different tasks, these coefficients can "jump" uncontrollably.

Excessively large fluctuations in the coefficients turn clear training signals into interference, destabilizing the entire system.

Until now, the standard tools have been GRPO and GSPO, which used the principle of hard clipping. If the coefficient exceeded the specified limits, the gradient was simply set to zero.

FIRST MINUS: Information loss. Valuable but outlier data were ruthlessly discarded.

SECOND MINUS: Impossible balance. If you make the limits too narrow, you stifle learning. If you make them too wide, parasitic noise creeps in. For capricious MoE architectures, this dilemma is particularly relevant.
SAPO proposes to abandon hard clipping in favor of intelligent smoothing.

Instead of abruptly setting the gradient to zero, SAPO uses a smooth, adaptive function (controlled by temperature) that gently reduces the influence of problematic gradients without completely nullifying them. This creates continuous confidence regions within which the model can learn more flexibly and safely.


Like GSPO, but smarter. If only one token in a long answer was wrong, GSPO punished the entire sequence. SAPO selectively suppresses only the "culprit", preserving useful signals from the rest of the words. This dramatically improves the efficiency of training data sets.

Like GRPO, but smoother. Instead of abruptly disabling the gradient for a bad token, SAPO applies a gradual attenuation. This prevents sudden jumps in learning and ensures a smooth and stable adjustment of the model's policy.

The icing on the cake of the method is an asymmetric temperature design. SAPO processes "good" and "bad" updates differently. For tokens with a negative contribution, a higher temperature is used, forcing their influence to attenuate faster and stronger.

This simple rule reliably suppresses the most dangerous fluctuations, which in practice leads to unprecedented stability of the RL learning process.

CONFIRMED BY TESTS

When training Qwen3-30B-A3B-Base, SAPO not only showed a more stable learning curve, but also achieved higher results on complex mathematical benchmarks AIME25, HMMT25. And he did this without the labor-intensive routing reproduction that competitors needed to work with MoE.

The success was repeated in a large-scale experiment with multimodal Qwen3-VL-30B-A3B, where SAPO consistently outperformed its analogues in mixed tasks on coding, logic, and mathematics.


#AI #ML #LLM #MoE #SAPO #Qwen

••••••••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
More from @dataxplore
  1. Oct 1, 2026Monitoring and debugging "silent" token drift in LLM pipelines: frequency analysis of embe…
  2. Sep 30, 2026Online detection of feature collisions in TDA transformation When using Topological Data A…
  3. Sep 18, 2026Am going to announce something big (for me, it's really big) on October 11, 2026.
  4. Sep 14, 2026Post #2188
  5. Aug 31, 2026I joined a Russian community on Telegram. They share some Russian startup and technology u…
  6. Aug 22, 2026Post #2185
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →