llm-d is a new Open Source project and community for scalable GenAI deployments in Kubernetes.
Described as “a Kubernetes-native high-performance distributed LLM inference framework,” llm-d leverages existing technologies, such as vLLM, Kubernetes, and Inference Gateway, to provide a vLLM-optimised inference scheduler, disaggregated serving, and disaggregated prefix caching. The authors also plan to implement a traffic- and hardware-aware autoscaler.
You can find more details on llm-d in:
- yesterday’s announcement by CoreWeave, Google, IBM Research, NVIDIA, and Red Hat;
- the project’s GitHub repo.
#news #aiml #tools
Post #209
864

- 👍 6