Landon Clipp built a GPU Containers as a Service platform from scratch — solving multi-tenant GPU isolation with Kata/QEMU, NVLink fabric partitioning, and Cilium network policies.
You will learn:
- Why standard NVIDIA tooling fails in multi-tenant setups, and how PCI topology scanning makes GPUs visible to Kubernetes without kernel drivers
- How to partition the NVLink fabric between tenants using a trusted service VM running Fabric Manager
- What caused 8-GPU VMs to take 30+ minutes to boot, and the fixes that brought it down to minutes
Watch (or listen to) it here: https://ku.bz/jjK_yJTDz
🌟 This episode is brought to you by LearnKube — get started on your Kubernetes journey through comprehensive online, in-person or remote training. https://learnkube.com/training
With @Birthmarkb
Post #1695
226
Forwarded from KubeFM