Landmark 2 stops to get here · leads to 1

Diffusion Model

A generative model that learns to denoise data, achieving state-of-the-art image generation (Stable Diffusion, DALL-E 2).

Your route here

2 stops · basics first
  1. Machine Learning ✓ understood

    Building systems that learn patterns from data instead of following hand-written rules, getting better at a task as they see more examples.

  2. Neural Network ✓ understood

    A computational model inspired by biological neural networks, consisting of interconnected nodes (neurons) organized in layers that process information through weighted connections.

  3. Diffusion Model · you are here ✓ understood

Picture it

Forward (noising)

  • Fixed, no learning
  • Add a little Gaussian noise
  • Repeat until pure noise
  • Makes training examples

Reverse (denoising)

  • Learned neural network
  • Predict and remove noise
  • Repeat step by step
  • Pure noise becomes an image
Notice that the model only learns the reverse direction; generation starts from random noise and denoises step by step.

Where it sits

Before this

Neural Network
Diffusion Model

Explore nearby

In the research

All papers →

20 papers that build on Diffusion Model ; showing 5, canon first.

Canon · 2020 Denoising Diffusion Probabilistic Models DDPM is the recipe behind Stable Diffusion, DALL·E, Midjourney and today’s video models. Canon · 2021 High-Resolution Image Synthesis with Latent Diffusion Models It’s the architecture behind Stable Diffusion, which put high-quality text-to-image generation on ordinary GPUs. Frontier · Aug 2025 · 875 citations Qwen-Image Technical Report Legible, editable text in generated images was a long-standing weakness; this open model largely fixes it. Frontier · Sep 2025 · 226 citations Seedream 4.0: Toward Next-generation Multimodal Image Generation It set the bar for commercial image generation and editing in one fast system. Frontier · Oct 2025 · 221 citations Diffusion Transformers with Representation Autoencoders It argues the generation and understanding sides of vision should share one latent space.