Want to generate a 5-second video on a large SOTA model? You Run Prompt→Go For Coffee → Come Back, process is still going on. Its Harsh reality that generation can take more than an hour.
🟢 How TurboDiffusion accelerats 100+ times?
The PROBLEM: Main problem were monstrous computational complexity of the attention mechanism in transformers, the need for hundreds of denoising steps, and the huge amount of memory for full-precision weights.
The SOLUTION: TurboDiffusion combined all effective compression and acceleration methods into a single pipeline. Idea was that sparsity and quantization are techniques that do not interfere with each other.
Architecture relies on 3 Optimization Pillars:
☞ They replaced the standard attention with a hybrid of SageAttention2++ and Sparse-Linear Attention (SLA), which turned quadratic complexity into linear. So that the model focuses only on important tokens.
☞ They distilled the sampling through rCM - instead of the standard 50-100 steps, the model comes to the result in just 3-4 steps without losing the essence of the image.
☞ They converted both the weights and activations of linear layers to INT8 using block quantization, so as not to lose accuracy.
To top it all off, they were able to combine the weights after fine-tuning under SLA and distillation of rCM into a single model, avoiding conflicts.
BENCHMARK results look like a typo, but it's not.
On the RTX 5090, the generation time for the heavy model Wan2.2-I2V 14B fell from 69 minutes to 35.4 seconds. And for the lighter Wan 2.1-1.3B - from almost 3 minutes to 1.8 seconds.
At the same time, judging by the examples, the visual quality remains practically indistinguishable from the original.
Set of Models, Technical report, GitHub • #AI #ML #I2V #T2V #TurboDiffusion
••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore