This is impossible : AWS network (EFA) does not support GPUDirect Async, so GPUs on different machines cannot exchange data fast enough.
🟢 How Perplexity find solution?
They built new software that transfers coordination to the CPU, allowing GPUs to still synchronize almost directly.
This makes inference of models with *1 trillion parameters* efficient on regular AWS clusters, not just on specialized supercomputers.
They prepared expert-parallel kernels for fast MoE inference on AWS EFA:
1T MoE works practically without degradation, and the multi-node mode is comparable to or faster than single-node on 671B DeepSeek V3 with medium batch sizes, opening the way to serving Kimi K2.
PROBLEM: EFA does not support GPUDirect Async, and the standard NVSHMEM-proxy provides MoE routing with latencies above 1 ms.
SOLUTION: kernels pack tokens into single RDMA writes directly from the GPU, and a special CPU thread launches the transfer and overlaps it with GEMM computations.
The result is EFA suddenly becomes a viable option for massive MoE inference.
This is solid engineering and a reasonable balance of accuracy and memory for teams needing portability between clouds.
🤖 Data Science, ML & Big Data with @DataXplore
