John McBride, VP of Infrastructure and AI Engineering at the Linux Foundation shares how using Kubernetes and open-source AI models saved them tens of thousands of dollars.
You will learn:
- How to deploy VLLM on Kubernetes to serve open-source LLMs like Mistral and Llama
- How running inference workloads on your own infrastructure with T4 GPUs can reduce costs from tens of thousands to just a couple thousand dollars monthly
- Practical approaches to monitoring GPU workloads in production, including handling unpredictable failures and VRAM consumption issues
Watch (or listen to) it here: https://ku.bz/wP6bTlrFs
🌟 This episode is brought to you by StackGen! Don't let infrastructure block your teams. StackGen deterministically generates secure cloud infrastructure from any input - existing cloud environments, IaC or application code https://ku.bz/t0gBX9qQz
With @Birthmarkb "SWAG expert" Farrell
Post #1294
372
Forwarded from KubeFM