Cursor open-sourcing Mixture-of-Kittens (MoK), their MoE training megakernel for NVL72s.
It fuses all MoE communication and computation into a single, fully deterministic kernel, and runs up to 2.37x faster than the strongest public baselines.
MoK now powers training across tens of thousands of GPUs at Cursor.
In production, it raised end-to-end training throughput by 1.41x over our previous DeepEP-based stack.
Post #4397
620