Officially release DeepSeek-V4.1-Flash: smarter, faster, more efficient.
The smallest model in a new architecture family, with native visual understanding.
Asymmetric architecture. More intelligence, less cost.
- 552B-parameter MoE.
- New Causal Encoder–Decoder architecture: just 8B active parameters for input, 16B for output.
- New pre-training methods + larger-scale RL post-training deliver benchmark results ahead of flagship models, including DeepSeek-V4-Pro.
Smaller KV cache. Bigger savings.
Paper.
Post #4465
637