An interesting post where the author explained that what we call text diffusion is actually just a generalized version of classic BERT training.
🟢 How does BERT work?
In BERT, the model takes text and masks some words, then learns to guess which words were hidden.
In diffusion, almost the same thing happens, but with more steps: at each step, the model slightly "damages" the text (adds noise), then restores it, losing less and less meaning until it collects the final clean text.
So BERT performs one denoising step, “GUESSING THE MASKED WORDS.”
And the diffusion model performs many such steps in a row, gradually turning a random set of tokens into meaningful text.
Barry fine-tuned RoBERTa to demonstrate this in practice and got a real text diffusion generator.
In the example:
- RoBERTa (an improved version of BERT) and the WikiText dataset are used.
- At each step, some tokens are replaced with <MASK>, the model restores them, then masks again — and so on several times.
- After several iterations, the model can generate coherent text, even without an autoregressive decoder (like GPT).
The author mentions that later he came across the work DiffusionBERT, where the idea was implemented more deeply and confirmed with results.
Main idea is that BERT can be considered a single-step version of text diffusion.
If you add more steps, you get a diffusion text generator.
The model generates meaningful text, although not perfectly coherent. If BERT is one diffusion step, then the future may belong to models combining "understanding" and "generation" of text in one process.
#AI #Diffusion #RoBERTa #BERT #LanguageModel #MLM #Research
🤖 Data Science, ML & Big Data with @DataXplore