Most often in the cloud, where several models are hosted, each model is tied to specific GPUs. For example, Llama-70B → 8× A100.
And even if no one is currently querying the model, the GPU remains reserved and idle because the weights are already loaded.
🟢 How Aegaeon helps?
Alibaba discovered that this seemingly innocent idling actually consumes a huge amount of resources. It turned out that in their cloud 17.7% of all GPUs were occupied by models that processed only 1.35% of all requests. First, this is terribly inefficient. Second, such a system is very hard to scale if more models appear.
So they took on optimization and proposed something called Aegaeon (don’t know how to pronounce it). This is a system where instead of a one-to-one mapping of "model-to-GPU", each GPU can handle multiple models simultaneously.
It’s somewhat similar to Kubernetes: the cluster turns into a unified pooling system that can dynamically allocate and free memory.
The main idea is that the system switches at the token level, not whole requests. Usually, a model is loaded into memory entirely and runs until it finishes the response. Aegaeon breaks the process into prefill and decode, alternating them between models right during generation.
This happens without full initialization: the scheduler caches the necessary parts in VRAM and loads the rest as needed. So there are delays, but minimal – within 3-5%.
Currently, Aegaeon is already running directly in Alibaba Cloud. And engineers claim they managed to reduce the number of required GPUs from 1192 to 213. That’s a minus 82%!
Necessity is the mother of invention who were banned from importing GPUs 🍿
🤖 Data Science, ML & Big Data with @DataXplore
