Allows for dramatically faster switching between language models.
🟢 How it is useful?
Traditional methods require either keeping both models loaded (which doubles GPU load) or reloading them sequentially with a 30–100 second pause. Sleep Mode offers a third option: models are "put to sleep" and "woken up" in seconds, preserving their already initialized state.
There are two sleep levels available:
Level 1️⃣ - Weights are offloaded to RAM, fast wake-up but requires a lot of RAM;
Level 2️⃣ - Weights are fully unloaded, minimal RAM usage, wake-up is slightly slower.
Both levels provide performance gains: model switching became 18 to 200 times faster, and inference time after waking up improved by 61–88%, since process memory, CUDA graphs, and JIT compilation are preserved.
Ideal for scenarios with frequent use of different models and makes multi-model serving practical even on mid-range GPUs - from A4000 to A100.
🟠 How to enable & test?
export VLLM_SERVER_DEV_MODE=1
vllm serve Qwen/Qwen3-0.6B --enable-sleep-mode --port 8000
# Sleep
curl -X POST 'localhost:8000/sleep?level=1'
#Wake
curl -X POST 'localhost:8000/wake_up'
# Level 2 only
curl -X POST 'localhost:8000/collective_rpc' \
-H 'Content-Type: application/json' \
-d '{"method":"reload_weights"}'
curl -X POST 'localhost:8000/reset_prefix_cache'
🤖 Data Science, ML & Big Data with @DataXplore
