🔴 What is the problem?
- MoE activates only a subset of experts per token → saving compute.
- But with large batch sizes, communication and memory increase:
- more experts are loaded,
- KV cache grows,
- memory and network become bottlenecks.
🟢 What is The solution ?
expert parallelism
- Experts are spread across many GPUs.
- Each token goes to top-N experts + a shared expert.
- In DeepSeek: 8 experts out of 256 per layer × 58 layers.
To handle communications:
- attention remains data parallel (cache stays on one GPU),
- only small activation vectors are communicated,
- two microbatches: one computes, the other communicates,
- hot experts are duplicated,
- tokens try to keep experts within a single node.
Optimizations
- multi-head latent attention → compresses KV cache to ~70KB instead of hundreds of KB.
- reworking attention math → fewer computations for long contexts.
- prefill and decode are separated, cache yields ~56% hits → less cost.
Economics
- Cost = $/GPU-hour ÷ tokens/hour.
- Cheaper with larger batch sizes, faster interconnects, more GPUs.
- But if the service promises 20 tokens/sec per user → smaller batches, higher price.
Practice
- NVLink clusters scale excellently.
- InfiniBand between DGX is a bottleneck.
- 72 GPUs at batch 64 → billions of tokens per day for about ~$0.40 / 1M tokens.
Inshort, MoEs become cheap with: Large batches, Compressed KV cache, Smart routing, Separation of prefill and decode, Fast interconnects.
This provides flexibility: fast chat sells for more, while bulk generation (synthetic data, fine-tuning) runs almost at cost.
🤖 Data Science, ML & Big Data with @DataXplore