TGViewer
Data eXplore : Data Science, ML, Big Data, LLMs and AI Security Data eXplore : Data Science, ML, Big Data, LLMs and AI Security @dataxplore · 585 subscribers
Post #1751 225
Alibaba found a way to reduce GPU demand by 82%

Most often in the cloud, where several models are hosted, each model is tied to specific GPUs. For example, Llama-70B → 8× A100.

And even if no one is currently querying the model, the GPU remains reserved and idle because the weights are already loaded.

🟢 How Aegaeon helps?
Alibaba discovered that this seemingly innocent idling actually consumes a huge amount of resources. It turned out that in their cloud 17.7% of all GPUs were occupied by models that processed only 1.35% of all requests. First, this is terribly inefficient. Second, such a system is very hard to scale if more models appear.

So they took on optimization and proposed something called Aegaeon (don’t know how to pronounce it). This is a system where instead of a one-to-one mapping of "model-to-GPU", each GPU can handle multiple models simultaneously.

It’s somewhat similar to Kubernetes: the cluster turns into a unified pooling system that can dynamically allocate and free memory.

The main idea is that the system switches at the token level, not whole requests. Usually, a model is loaded into memory entirely and runs until it finishes the response. Aegaeon breaks the process into prefill and decode, alternating them between models right during generation.

This happens without full initialization: the scheduler caches the necessary parts in VRAM and loads the rest as needed. So there are delays, but minimal – within 3-5%.

Currently, Aegaeon is already running directly in Alibaba Cloud. And engineers claim they managed to reduce the number of required GPUs from 1192 to 213. That’s a minus 82%!


Necessity is the mother of invention who were banned from importing GPUs 🍿

🤖 Data Science, ML & Big Data with @DataXplore
More from @dataxplore
  1. Oct 1, 2026Monitoring and debugging "silent" token drift in LLM pipelines: frequency analysis of embe…
  2. Sep 30, 2026Online detection of feature collisions in TDA transformation When using Topological Data A…
  3. Sep 18, 2026Am going to announce something big (for me, it's really big) on October 11, 2026.
  4. Sep 14, 2026Post #2188
  5. Aug 31, 2026I joined a Russian community on Telegram. They share some Russian startup and technology u…
  6. Aug 22, 2026Post #2185
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →