TGViewer
Artificial Intelligence AI News Artificial Intelligence AI News @machinelearningresearchnews · 3.5K subscribers
Post #1547 875
Sparse attention already cut long-context compute. The KV cache sitting in HBM and on SSD is now the bottleneck DeepSeek AI went after, and they cut theirs to 890 bytes per token.

They released DeepSeek-V4.1-Flash, a 552B MoE model with 1M-token context that activates only 8B parameters per token during prefill and 16B during decode. The global KV cache footprint is roughly 1/4 of DeepSeek-V4-Flash and about 437x smaller than DeepSeek-V1, and the weights are open under an MIT license.

Here's what's actually interesting:


→ Causal Encoder-Decoder: prompt tokens stop at layer 20, decoder global KV is projected from the encoder's final hidden state, prefill compute nearly halves

→ Compressed Sparse Attention 2: three statically assigned layer modes, Reindex reuses main KV and indexer K, Reuse also reuses Top-K indices

→ FP4 main KV cache via quantization-aware training, SWA Bounded Replay replays only 128 tokens instead of layers times window

→ Terminal-Bench 2.1: 90.6 vs 89.1 for Opus-5

→ DeepSWE v1.1: 74.2 vs 74.0 for Opus-5 and 73.0 for GPT-5.6 Sol

→ Terminal-Bench 4.0: 31.2 vs 51.8 for Opus-5, so the gap on expert-level science tasks is still real


Full analysis: https://www.marktechpost.com/2026/09/10/deepseek-ai-released-deepseek-v4-1-flash-with-1m-context-fp4-kv-cache-and-cross-layer-attention-reuse/

Technical Details: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash

Paper: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/main/DeepSeek_V41_Tech_Report.pdf
  • 🤯 2
  • 🔥 1
More from @machinelearningresearchnews
  1. Oct 10, 2026Microsoft just released Microsoft-Decision-1, a decision-scoring model that returns calibr…
  2. Oct 7, 2026Meta just open-sourced Rebalancer, the assignment solver that has run resource allocation…
  3. Oct 6, 2026Mistral AI Releases Mistral Large 4 (Le Chonk): A 1.05T Parameter Multimodal MoE Model Mis…
  4. Oct 5, 2026[AI Model Family Series #1] We just published the complete story of Alibaba's Qwen: every…
  5. Oct 2, 2026Cloudflare Releases Clef: Open-Weight Decision Models That Return Typed Probabilities Inst…
  6. Sep 30, 2026Google DeepMind just announced Gemini 4 Argon, raising the output limit from 64K to 1M tok…
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →