Fast-dLLM v2: Efficient Block-Diffusion LLM
Chengyue Wu et al.
arXiv:2509.26328
In short
Fast-dLLM v2 converts a pretrained autoregressive LLM into a block-diffusion model that generates chunks of tokens in parallel, using only about a billion tokens of fine-tuning. Hierarchical caching and parallel decoding give up to 2.5× speed-ups without losing accuracy.
Why it matters
A cheap path to faster inference from models you already have.
Read first
The 4 Field Guide ideas this paper leans on.
Starting from scratch? The full route 16 ideas · basics first
- Natural Language Processing ✓ understood
The field of AI that lets computers read, interpret, translate and generate human language, from spam filters and search to chatbots.
- Token ✓ understood
The basic unit of text that a language model processes, typically representing a word, subword, or character. Tokens are the fundamental building blocks for LLM input and output.
- Tokenization ✓ understood
Splitting text into tokens, usually subword pieces, and mapping each to an integer ID so a language model can process it.
- Language Modeling ✓ understood
Learning probability distributions over sequences of words to predict what comes next.
- Autoregressive Model · read first ✓ understood
A model that generates output one token at a time, using previously generated tokens as input for the next prediction.
- Machine Learning ✓ understood
Building systems that learn patterns from data instead of following hand-written rules, getting better at a task as they see more examples.
- Neural Network ✓ understood
A computational model inspired by biological neural networks, consisting of interconnected nodes (neurons) organized in layers that process information through weighted connections.
- Diffusion Model · read first ✓ understood
A generative model that learns to denoise data, achieving state-of-the-art image generation (Stable Diffusion, DALL-E 2).
- Dataset ✓ understood
A collection of data examples used for training, validating, or testing machine learning models.
- Training ✓ understood
The process of fitting a model to data by repeatedly measuring how wrong its outputs are and adjusting its parameters to reduce that error.
- Inference ✓ understood
Running a trained model on new inputs to get predictions, with its weights frozen: the stage of a model's life that users actually interact with.
- Inference Latency · read first ✓ understood
The time delay between submitting input and receiving output from a deployed model, critical for real-time applications.
- Unsupervised Learning ✓ understood
Learning from unlabeled data to discover hidden patterns, structures, or relationships without explicit target outputs.
- Self-Supervised Learning ✓ understood
Learning representations from unlabeled data by creating supervised tasks from the data itself (masked prediction, contrastive learning).
- Pre-training ✓ understood
Training a model on a large dataset (often self-supervised) before fine-tuning on specific tasks, enabling transfer learning.
- Fine-Tuning · read first ✓ understood
The process of further training a pre-trained model on a specific dataset to adapt it for a particular task or domain.
In the frontier
- Rank
- #44 of 100
- Citations
- 114
- as of Aug 9, 2026
- Published
- Sep 2025
Topics: Diffusion LMs and decoding , Efficiency and serving , Model architecture
Selection: 1kpapers.com by Together AI, most-cited as of Aug 9, 2026