TGViewer
无印🐑品 #BeHonest 无印🐑品 #BeHonest @beetselect · 1.76K subscribers
Post #18573 2.15K
DeepSeek V4.1 Flash 为 552B 参数的 MoE 模型,采用了全新的 Causal-Encoder-Decoder 结构,输入和输出不对称,输入激活只有 8B,输出激活 16B,成本显著低于已知的同尺寸模型。同时,V4.1 Flash 还采用了新的预训练方式、经过了更大规模的强化学习后训练,在基准测试中,成功超越了包括 DeepSeek V4 Pro 在内的一众旗舰模型的智能水平。

新一代模型大幅减少了 KV Cache 缓存的大小,与上一代模型相比,对 HBM 的需求减少到 1/4,对 SSD 的需求减少到 1/8。在 Agent 使用场景中,缓存命中的费用往往占比较高,对 KV Cache 的压缩大幅降低了 Agent 类任务的使用成本。

https://mp.weixin.qq.com/s/qg0NU3NNUbp1co2PdkAPAg
More from @beetselect
  1. Oct 1, 2026西恩西,西恩西
  2. Oct 1, 2026? https://www.ithome.com/1/009/081.htm
  3. Oct 1, 2026无印🐑品 #BeHonest pinned a photo
  4. Oct 1, 2026photo post
  5. Oct 1, 2026华为黑卡来了
  6. Oct 1, 2026所以还是有 9050 non-Pro 的,看看谁用
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →