Remember Tetris. You need to fit the pieces as tightly as possible so there are no empty spaces. A very similar problem arises in cloud data centers like Google Cloud.
🟢 In-depth Explanation!?
There are physical servers running virtual machines for different tasks. These VMs appear, run for some time, and then disappear. Some VMs are taken for testing and run for 15-20 minutes, while others host databases for months.
You cannot know in advance how long a VM will live. At the same time, there is a specific optimization problem: to pack them in a way that uses resources as efficiently and densely as possible. Just like in Tetris.
Simple optimization does not work here precisely because of the uncertainty. So Google thought it through and attached a probabilistic ML model.
It predicts the probability distribution of a VM's lifetime based on the general distribution (which, by the way, is heavily skewed), VM metadata, user behavior, creation method, etc. The output is something like "With 80% probability this VM will live for an hour, with 15% probability – a day, and 5% – longer than a week." This is called survival analysis.
Interestingly, the forecast is dynamic and updates over time. For example, if the VM is still running after 10 days, the model revises the estimate.
And based on this predicted distribution, optimization algorithms operate. For example, a scheduler tries to place several identical VMs on one server to free it completely later. Or an algorithm places short-lived VMs on servers with long-lived ones to fill small gaps that would otherwise be lost.
And the metrics. Google has already tested this approach on their servers and (attention!) equipment downtime has decreased on average by 5%. Imagine how much that is in dollars
Great work and a cool case 🙂
🤖 Data Science, ML & Big Data with @DataXplore
