🟢 All the How's?
• DeepSeek-V3.2-Speciale outperforms Gemini 3.0 Pro in mathematics and code
• The new flagship model combines reasoning + agency
• Architecture MoE from the V3.1 Terminus family, context 128k
• The main innovation - DeepSeek Sparse Attention (DSA), designed for cheap long context
What makes DSA
Normal attention - O(T²), which is costly at 128k tokens.
DSA reduces the cost to O(T·U), where U is only a small number of relevant tokens.
How it works:
1) Lightning Indexer - a lightweight network evaluates the importance of each previous token
2) Fine-grained top-k - the model selects only the most useful tokens and calculates attention on them
How it was trained
We started with a checkpoint of V3.1 (128k) and performed two-stage fine-tuning:
• Stage 1 - dense attention, frozen model, trained only DSA
• Stage 2 - gradual transition to DSA across the entire model
Result: long context has become really cheap, and the quality is higher than previous versions and competitors.
••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore