TGViewer
Data eXplore : Data Science, ML, Big Data, LLMs and AI Security Data eXplore : Data Science, ML, Big Data, LLMs and AI Security @dataxplore · 585 subscribers
Post #1733 114
How Mixture-of-Experts (MoE) LLMs can be made truly cheap by tailoring the architecture to the hardware? (analysis)

🔴 What is the problem?
- MoE activates only a subset of experts per token → saving compute.
- But with large batch sizes, communication and memory increase:
- more experts are loaded,
- KV cache grows,
- memory and network become bottlenecks.


🟢 What is The solution ?
expert parallelism
- Experts are spread across many GPUs.
- Each token goes to top-N experts + a shared expert.
- In DeepSeek: 8 experts out of 256 per layer × 58 layers.

To handle communications:
- attention remains data parallel (cache stays on one GPU),
- only small activation vectors are communicated,
- two microbatches: one computes, the other communicates,
- hot experts are duplicated,
- tokens try to keep experts within a single node.

Optimizations
- multi-head latent attention → compresses KV cache to ~70KB instead of hundreds of KB.
- reworking attention math → fewer computations for long contexts.
- prefill and decode are separated, cache yields ~56% hits → less cost.

Economics
- Cost = $/GPU-hour ÷ tokens/hour.
- Cheaper with larger batch sizes, faster interconnects, more GPUs.
- But if the service promises 20 tokens/sec per user → smaller batches, higher price.

Practice
- NVLink clusters scale excellently.
- InfiniBand between DGX is a bottleneck.
- 72 GPUs at batch 64 → billions of tokens per day for about ~$0.40 / 1M tokens.


Inshort, MoEs become cheap with: Large batches, Compressed KV cache, Smart routing, Separation of prefill and decode, Fast interconnects.

This provides flexibility: fast chat sells for more, while bulk generation (synthetic data, fine-tuning) runs almost at cost.

🤖 Data Science, ML & Big Data with @DataXplore
More from @dataxplore
  1. Oct 1, 2026Monitoring and debugging "silent" token drift in LLM pipelines: frequency analysis of embe…
  2. Sep 30, 2026Online detection of feature collisions in TDA transformation When using Topological Data A…
  3. Sep 18, 2026Am going to announce something big (for me, it's really big) on October 11, 2026.
  4. Sep 14, 2026Post #2188
  5. Aug 31, 2026I joined a Russian community on Telegram. They share some Russian startup and technology u…
  6. Aug 22, 2026Post #2185
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →