Bigger MoE models keep winning on quality, but serving them at interactive latency is still hard.
NVIDIA compresses the hybrid MoE Nemotron-3-Super into Puzzle-75B-A9B and roughly doubles interactive server throughput while holding quality.
Paper: https://arxiv.org/abs/2607.04371
Post #1651
5.76K
