Новости Линукс Linux
По всем вопросам @evgenycarter
Post #19954
65

Qwen 3.8 27B exposes the RTX 5090’s inference-engine bottleneck
Qwen 3.8 27B can fit on a single RTX 5090, but real-world performance depends more on the inference engine and context configuration than on the card’s 32 GB of VRAM. The same GPU can take roughly 30 minutes to produce a long-context response with llama.cpp, deliver about 20 tokens per second through a basic vLLM setup, or reach around 200 tokens per second with a more optimized engine.
Source
👉@sysadminoff
https://4sysops.com/archives/qwen-3-8-27b-exposes-the-rtx-5090s-inference-engine-bottleneck/
Qwen 3.8 27B can fit on a single RTX 5090, but real-world performance depends more on the inference engine and context configuration than on the card’s 32 GB of VRAM. The same GPU can take roughly 30 minutes to produce a long-context response with llama.cpp, deliver about 20 tokens per second through a basic vLLM setup, or reach around 200 tokens per second with a more optimized engine.
Source
👉@sysadminoff
https://4sysops.com/archives/qwen-3-8-27b-exposes-the-rtx-5090s-inference-engine-bottleneck/






