Sparse attention already cut long-context compute. The KV cache sitting in HBM and on SSD is now the bottleneck DeepSeek AI went after, and they cut theirs to 890 bytes per token.
They released DeepSeek-V4.1-Flash, a 552B MoE model with 1M-token context that activates only 8B parameters per token during prefill and 16B during decode. The global KV cache footprint is roughly 1/4 of DeepSeek-V4-Flash and about 437x smaller than DeepSeek-V1, and the weights are open under an MIT license.
Here's what's actually interesting:
→ Causal Encoder-Decoder: prompt tokens stop at layer 20, decoder global KV is projected from the encoder's final hidden state, prefill compute nearly halves
→ Compressed Sparse Attention 2: three statically assigned layer modes, Reindex reuses main KV and indexer K, Reuse also reuses Top-K indices
→ FP4 main KV cache via quantization-aware training, SWA Bounded Replay replays only 128 tokens instead of layers times window
→ Terminal-Bench 2.1: 90.6 vs 89.1 for Opus-5
→ DeepSWE v1.1: 74.2 vs 74.0 for Opus-5 and 73.0 for GPT-5.6 Sol
→ Terminal-Bench 4.0: 31.2 vs 51.8 for Opus-5, so the gap on expert-level science tasks is still real
Full analysis: https://www.marktechpost.com/2026/09/10/deepseek-ai-released-deepseek-v4-1-flash-with-1m-context-fp4-kv-cache-and-cross-layer-attention-reuse/
Technical Details: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash
Paper: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/main/DeepSeek_V41_Tech_Report.pdf
Post #1547
875

- 🤯 2
- 🔥 1