Fact-checked Aug 10, 2026
Also called: diffusion models, DDPM, denoising diffusion probabilistic models
A diffusion model is a type of generative artificial intelligence that learns to create new data, such as images, audio, or video, by gradually refining random noise into a clear output.
Diffusion models are a powerful class of generative AI. Generative AI refers to artificial intelligence systems that can produce new content, rather than just analyzing existing content. Diffusion models are particularly good at tasks like turning text descriptions into realistic images, or creating unique audio clips and even short videos. They've become incredibly popular for their ability to generate high-quality, diverse, and creative outputs.
Imagine you have a clear photograph, and you slowly add random visual static, or "noise," to it, step by step, until the original image is completely obscured and all you see is static. A diffusion model works by learning to reverse this process. During its training, the model is shown many images and learns how to systematically remove the noise, one tiny step at a time, to recover the original image from a noisy version. It becomes very good at predicting what information needs to be added back to make the picture clearer.
When a diffusion model is asked to generate something new, it starts with a canvas of pure random noise. Think of it like a blank slate of TV static. Then, using the knowledge it gained during training, the model iteratively "denoises" this static, slowly transforming it into a coherent image, sound, or other data type. Each step makes the output a little bit clearer, guided by whatever prompt or condition it was given, until a complete and recognizable piece of content emerges.
A great example of where you'd encounter diffusion models is in popular text-to-image generators like Stable Diffusion or DALL-E. When you type in a prompt like "a cat wearing a spacesuit," the diffusion model takes this instruction, starts with a random noisy image, and progressively refines that noise until it creates an image matching your description. These models are also used in areas like video editing, creating synthetic data for training other AIs, and even designing new materials.
One common misconception is that these models truly "understand" the world or the concepts they generate. In reality, they are sophisticated pattern-matching systems. They learn statistical relationships between text and pixels, or between different patterns of noise and clear data. While they can create incredibly realistic and creative outputs, they don't possess conscious understanding or common-sense reasoning, which can sometimes lead to strange or nonsensical results, especially when dealing with complex scenes or specific anatomical details like hands.
Diffusion models are a powerful class of generative AI. Generative AI refers to artificial intelligence systems that can produce new content, rather than just analyzing existing content. Diffusion models are particularly good at tasks like turning text descriptions into realistic images, or creating unique audio clips and even short videos. They've become incredibly popular for their ability to generate high-quality, diverse, and creative outputs.
diffusion model is also referred to as diffusion models, DDPM, denoising diffusion probabilistic models.
Daily Deck explains terms like diffusion model as part of a free seven-card daily brief. No jargon. No fluff.
Start free