Federico Iezzi, Customer Engineer at Google Cloud, explains how his team achieved 1 million output tokens per second using Qwen 3.5 27B, vLLM, GKE Autopilot, and NVIDIA B200 GPUs.
You will learn:
- Why memory bandwidth limits decode performance
- How Federico chose between tensor and data parallelism
- What changed after enabling multi-token prediction and reducing the KV cache footprint with FP8 quantization
Watch (or listen to) it here: https://ku.bz/1xD9Md0mb
🌟 This episode is brought to you by LearnKube. Download the free book, The Technical Guide to Kubernetes Rightsizing, to understand what Prometheus and Grafana cannot tell you about safely reducing requests and limits: https://learnkube.com/kubernetes-rightsizing
With @Birthmarkb
Post #1834
213
Forwarded from KubeFM