TGViewer
Data eXplore : Data Science, ML, Big Data, LLMs and AI Security Data eXplore : Data Science, ML, Big Data, LLMs and AI Security @dataxplore · 583 subscribers
Post #1920 191
Hybrid tokenizer for diffusion on steroids.

In diffusion architectures, VAE is a thankless task which is convert pixels into latent code and back, and adding data to it doesn't help main DiT model generate better images.

🟢 What Visual Tokenizer Pre-training (VTP) Changed?

Their hypothesis is that the tokenizer should not just mechanically "zip up" pixels, but understand the semantics of the image.

To implement this, they combined 3 losses in the training of the tokenizer:

☞ Standard pixel reconstruction loss;
☞ Self-supervised learning (through Masked Image Modeling and distillation, as in DINOv2);
☞ Image-text contrastive loss (as in CLIP).

This made the latent space structurally semantic: now the vectors encode meanings, not just color spots.

Theoretical assumptions were confirmed in practice.

It turned out that the quality of generation directly depends on the "intelligence" of the tokenizer. Without changing the architecture and hyperparameters of DiT itself and without increasing the cost of its training, just by using the VTP tokenizer, it was possible to improve the FID metric by 65.8% and accelerate the convergence of the model by 3 times.

But the main discovery is that the scaling law for Stage 1 has worked.

Now, the more computational power and data are poured into the pre-training of the tokenizer, the better the final generation becomes, which was impossible to achieve with ordinary VAEs before.

3 VTP checkpoints with different numbers of parameters have been published in the open access: VTP-Large, VTP-Base, VTP-Small.


TileSet of models with GitHub • #AI #ML #Diffusion #Tokenizer #Minimax

••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
More from @dataxplore
  1. Oct 1, 2026Monitoring and debugging "silent" token drift in LLM pipelines: frequency analysis of embe…
  2. Sep 30, 2026Online detection of feature collisions in TDA transformation When using Topological Data A…
  3. Sep 18, 2026Am going to announce something big (for me, it's really big) on October 11, 2026.
  4. Sep 14, 2026Post #2188
  5. Aug 31, 2026I joined a Russian community on Telegram. They share some Russian startup and technology u…
  6. Aug 22, 2026Post #2185
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →