๐จ Computer Vision ยท Generative AI
How Diffusion Models Work:
From Pure Noise to Photorealistic Image
At 2am, a game developer in Warsaw typed "a lighthouse in a thunderstorm, oil painting" into Midjourney. Four seconds later she had an image a concept artist would have quoted $400 for. She didn't think about the maths. But what happened in those four seconds is one of the most elegant ideas in modern machine learning: a model that learned to see by learning to un-see.
Every image you've ever seen from Midjourney, DALL-E, Stable Diffusion, or Flux was made the same way: start with pure random noise โ television static โ and remove the noise in small learned steps until an image emerges. That's the whole trick. Everything else in this guide is the detail that makes the trick work: the forward process that teaches the model to forget, the reverse process that teaches it to see, the text encoder that steers the noise, and the samplers that cut 1,000 steps down to 20.
๐ก Core idea: A diffusion model learns to reverse a noise-adding process. Training corrupts real images into pure static; inference runs the film backwards, removing predicted noise step by step until a meaningful image appears.
The Forward Process: Teaching a Model to Forget
During training, the model runs the forward diffusion process: take a real image and add a tiny amount of Gaussian noise at each of many timesteps โ classically 1,000 of them. After enough steps, the original image is statistically indistinguishable from pure noise. A photo of your dog, a satellite image, a painting โ all of them end up as the same featureless static.
That destruction is deliberately simple, and that's the point. Because the noise schedule is known and controlled, the model always knows exactly how much noise was added at each step. The destruction is the lesson plan; the model's entire job is to learn the reverse.
The Reverse Process: Learning to See
At inference, the model runs the film backwards. It starts with pure random noise and asks its neural network โ a U-Net in the classic architectures, a transformer in newer ones like Flux โ one question per step: "what noise is present in this image right now?" Subtract the predicted noise and you get a slightly cleaner image. Repeat, and static becomes texture, texture becomes shapes, shapes become a lighthouse in a thunderstorm.
๐ฏ The network's job, precisely: at each timestep the network takes the noisy image and predicts the noise component added at that step. Remove the prediction, step backwards, repeat. The model never "draws" โ it only ever subtracts noise it recognises.
How Text Prompts Steer the Noise
Text-to-image works because the denoising is conditioned. A text encoder โ CLIP in Stable Diffusion, T5 in Flux and Imagen โ turns your prompt into embeddings that are fed into the network at every denoising step, pulling the noise removal toward images that match the description. Classifier-free guidance is the dial: at guidance 2 you get dreamy, loose interpretations; at 12 you get literal, sometimes brittle adherence.
Two details matter more than most users realise. First, the conditioning happens through the attention mechanism โ the same cross-attention that lets transformers weigh which words matter โ which is why "a red cube to the left of a blue sphere" actually places objects correctly. Second, the text encoder only sees a short window of your prompt โ about 77 tokens in CLIP โ which is why the back half of a long prompt gets silently ignored; our explainer on context windows covers why token limits shape what any model can "read."
The Model Families Powering 2026
| Family | Maker / Licence | Known For | Runs |
|---|---|---|---|
| Flux | Black Forest Labs, open weights | Text rendering, prompt adherence | Local, 12-24GB VRAM |
| Stable Diffusion 3.5 / SDXL | Stability AI, open weights | Huge LoRA / ControlNet ecosystem | Local, 8-12GB VRAM |
| Midjourney v7 | Midjourney, closed | Aesthetic quality, style range | Web/Discord, $10-60/mo |
| DALL-E 3 | OpenAI, closed | Prompt adherence via ChatGPT rewriting | API |
| Imagen 3 | Google, closed | Photorealism, clean typography | API |
Samplers: Why Your Image Takes 20 Steps, Not 1,000
The original DDPM paper needed around 1,000 tiny steps per image โ minutes of compute. Modern schedulers take fewer, smarter steps, and distilled models take almost none. Here's the practical landscape:
| Sampler | Typical Steps | Character | Use When |
|---|---|---|---|
| DDPM | ~1,000 | Research baseline, slow | Papers, not production |
| DDIM | ~50 | Deterministic, reproducible | Consistent iterations |
| DPM++ 2M Karras | 20-30 | Clean, reliable default | Photorealism, products |
| Euler a | 20-30 | Ancestral, more creative variance | Concept art, ideation |
| Turbo / Lightning / schnell | 1-8 | Distilled, near-instant | Real-time previews |
The practical rule I use: DPM++ 2M Karras at 25-30 steps for anything that needs to look real; Euler a when I want surprise; a distilled model when a human is waiting on the other side of the render.
Why Diffusion Beat GANs
- Stable training: no adversarial min-max game to collapse โ just a well-posed denoising objective.
- Diversity: each seed produces a different plausible image; GANs famously mode-collapse to a few favourites.
- Scalability: more data and compute reliably improve quality; GANs get unstable at scale.
- Editability: inpainting, img2img, and ControlNet fall out of the same denoising math almost for free.
What Goes Wrong: Hands, Text, and Copyright
Diffusion models fail in characteristic ways. Hands and fine geometry break because the model learns local texture statistics better than global structure. Text rendering failed for years until Flux and Imagen 3 largely fixed it with better encoders. And the legal layer is still genuinely unsettled: training-data lawsuits against model makers are ongoing, US copyright office guidance currently denies copyright to purely machine-generated images, and the EU AI Act adds transparency and labelling obligations.
The deeper question underneath all of it โ how do we make these models do what we actually want, not what we typed โ is the alignment problem, and it applies to image models as much as to chatbots. A model that renders "no cars" as a car park full of cars is misaligned in the same technical sense as a chatbot that games its reward.
โ ๏ธ If you ship generated images commercially: check the model's licence (Midjourney paid tiers, SD/Flux open licences), keep provenance metadata (C2PA) where possible, and disclose generation where regulators or platforms require it. The legal floor is still being written โ build so you comply with the strictest version likely to pass.
Getting Better Outputs: A Practical Checklist
Before You Blame the Model:
- Be specific about the camera: "35mm, f/1.8, golden hour" changes more than any adjective.
- Front-load the prompt: the first ~20 tokens carry most of the weight in CLIP-based models.
- Fix the seed when iterating a composition; change one variable at a time.
- Use negative prompts for the artefacts you keep seeing ("extra fingers, watermark").
- Use img2img / inpainting to fix a region instead of re-rolling the whole image.
- Use ControlNet / reference images when composition matters more than surprise.
Where This Is Heading
The same denoising maths that renders a lighthouse now renders video โ Sora, Veo, and open video models are diffusion over spacetime, not just space. And video models are becoming world simulators: models that predict what happens next, which is the perceptual substrate that embodied AI needs. The path from text-to-image to robots that understand a kitchen is shorter than it looks; our coverage of LLMs in physical space traces exactly that line.
Whether any of it adds up to something bigger โ whether stacking these perceptual models gets us to general intelligence on the timelines people quote โ is the question we examine in AGI by 2027? A measured look at the evidence. What's not in doubt is the present tense: every product shot, game texture, and storyboard frame you'll see this year is, underneath, a neural network subtracting noise it was taught to recognise.