Looking for a way to simplify deploying LLMs on Kubernetes? This project provides everything you might need.
llmaz is an inference platform that integrates various Open Source projects for running LLMs. It supports:
- Different inference backends: vLLM, llama.cpp, Ollama, Text-Generation-Inference, SGLang, and TensorRT-LLM.
- Different model providers: HuggingFace, ModelScope, and ObjectStores.
- Chatbot interface based on Open WebUI.
- Heterogeneous devices.
- Distributed inference via multi-host and homogeneous xPyD support with LeaderWorkerSet.
- Envoy AI Gateway for token-based rate limiting, model routing, and more.
- Horizontal Pod scaling (HPA) and node autoscaling (Karpenter).
▶️ GitHub repo
Language: Go | License: Apache 2.0 | 278 ⭐️
#tools #aiml
Post #304
1.4K
- ❤ 3
- 👍 2