IDLM: Inverse-distilled Diffusion Language Models (ICML 2026, Our recent work)
Paper | Code | Checkpoints
⚡️ Can a language model generate a 1024-token sequence in just 16 forward passes?
That would mean producing 1024/16=64 tokens per forward pass.
For today’s standard language models, this sounds almost impossible. They are autoregressive, meaning they generate text token by token: first token, then the next, then the next…
So generating 1024 tokens usually requires 1024 forward passes.
This is one of the biggest bottlenecks in LLM inference.
A promising alternative is Diffusion Language Models. Instead of generating tokens one by one, they try to generate or refine the whole sequence in parallel, potentially removing the need for strict autoregressive decoding.
In theory, this could make generation much faster.
But in practice, diffusion-based language models often turn out to be slower, not faster, than autoregressive models.
The main challenge is the space dimension.
If we want the model to generate the next 64 tokens at once, it is not enough to predict one token 64 times independently. Ideally, the model should approximate the joint distribution over all possible 64-token continuations.
But the number of such continuations is enormous.
Even for a relatively small vocabulary, like GPT-2’s vocabulary of about ≈60,000 tokens, the number of possible 64-token sequences is:
60,000⁶⁴ ≈ 10³⁰⁶
That is an astronomically large number. We cannot enumerate these possibilities, store them, or explicitly simulate such a distribution.
So the real question becomes:
Can we generate many tokens in parallel while keeping the model’s complexity only linear in sequence length?
Post #714
1.86K
