TGViewer
DeepSeek DeepSeek @deepseek_ai · 1.42K subscribers
Post #87 820
💾 Smaller KV cache. Bigger savings.

Compared with the previous generation, V4.1-Flash’s KV cache needs just:
🔹 1/4 the HBM
🔹 1/8 the SSD storage
Cache-hit charges often account for a large share of agent costs. Compressing the cache cuts those costs significantly.
  • 🐳 11
  • 🤡 4
More from @deepseek_ai
  1. Sep 10, 2026🌐 Supporting open source. Expanding deployment options. We’ll work closely with the open-…
  2. Sep 10, 2026💰 More efficient architecture. Lower API prices. V4.1-Flash lets us serve more users at a…
  3. Sep 10, 2026⚡ V4.1-Flash is now live on the DeepSeek API with native multimodal support. Set your mode…
  4. Sep 10, 2026🧠 Asymmetric architecture. More intelligence, less cost. 🔹 552B-parameter MoE. 🔹 New Ca…
  5. Sep 10, 2026🚀 Introducing DeepSeek-V4.1-Flash: smarter, faster, more efficient. 🔹 Introducing the sm…
  6. Aug 21, 2026Files API is now live. 📁 🔹 Free to use 🔹 Upload an image once, then reference it by fil…
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →