In diffusion architectures, VAE is a thankless task which is convert pixels into latent code and back, and adding data to it doesn't help main DiT model generate better images.
🟢 What Visual Tokenizer Pre-training (VTP) Changed?
Their hypothesis is that the tokenizer should not just mechanically "zip up" pixels, but understand the semantics of the image.
To implement this, they combined 3 losses in the training of the tokenizer:
☞ Standard pixel reconstruction loss;
☞ Self-supervised learning (through Masked Image Modeling and distillation, as in DINOv2);
☞ Image-text contrastive loss (as in CLIP).
This made the latent space structurally semantic: now the vectors encode meanings, not just color spots.
Theoretical assumptions were confirmed in practice.
It turned out that the quality of generation directly depends on the "intelligence" of the tokenizer. Without changing the architecture and hyperparameters of DiT itself and without increasing the cost of its training, just by using the VTP tokenizer, it was possible to improve the FID metric by 65.8% and accelerate the convergence of the model by 3 times.
But the main discovery is that the scaling law for Stage 1 has worked.
Now, the more computational power and data are poured into the pre-training of the tokenizer, the better the final generation becomes, which was impossible to achieve with ordinary VAEs before.
3 VTP checkpoints with different numbers of parameters have been published in the open access: VTP-Large, VTP-Base, VTP-Small.
TileSet of models with GitHub • #AI #ML #Diffusion #Tokenizer #Minimax
••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
