The short answer
Quick answer: Most AI image generators are diffusion models. During training, the model is shown real images with increasing amounts of random noise added, and it learns to predict and remove that noise. To create a picture, the process runs in reverse: start from pure random noise and denoise it step by step until an image appears. Your text prompt is converted into numbers by a text encoder, and those numbers steer every denoising step so the emerging image matches the description. To keep this fast, the work is done on a small compressed version of the image, in latent space, which is expanded to full resolution at the end.
The core idea: learning to remove noise
Picture a photograph. Add a little random static. Add more. Keep going, and after enough steps the photograph is indistinguishable from television snow.
That destruction is easy and needs no intelligence. The reverse is the hard part: given a noisy image, work out what noise was added so it can be subtracted.
A diffusion model is a neural network trained to do exactly that. The approach was established in the 2020 paper Denoising Diffusion Probabilistic Models.
Training
- Take a real image from the training set.
- Pick a random noise level and add that much noise.
- Ask the network: "what noise was added?"
- Compare its guess with the truth and adjust the weights. See how neural networks learn.
- Repeat hundreds of millions of times, across all noise levels.
To predict the noise in a blurry picture of a dog, the network has to know something about what dogs look like. By learning to denoise, it learns the structure of images in general.
Generating
- Start with pure random noise.
- Ask the network to predict the noise, and remove a portion of it.
- Repeat, typically for 20 to 50 steps.
Structure emerges gradually: rough blobs of colour and composition first, then shapes, then fine detail. A different random starting point gives a different image. That starting point is set by the seed, which is why reusing a seed with the same prompt and settings reproduces a picture.
How text steers the picture
An unguided diffusion model would produce random plausible images. The prompt must direct it.
Text encoder. The prompt is turned into a set of vectors that capture its meaning. Many systems use encoders derived from CLIP, introduced in Learning Transferable Visual Models From Natural Language Supervision. CLIP was trained on hundreds of millions of image and caption pairs so that a picture and its description are mapped close together. It is a shared embedding space for text and images. Newer systems often use larger language-model encoders, which follow long, detailed prompts better.
Cross-attention. Inside the denoising network, the image being formed "looks at" the prompt's vectors at each step, using attention. This is how the word "red" ends up influencing the region where the car is.
Guidance. A technique called classifier-free guidance makes the image follow the prompt more closely. At each step the model predicts the noise twice, once with the prompt and once without, and then exaggerates the difference between them. The guidance scale setting controls how much:
- Low: looser, more varied, may ignore parts of the prompt.
- High: sticks to the prompt, but can look oversaturated or artificial.
A negative prompt uses the same mechanism to push the image away from things you do not want.
Latent space: the trick that made it practical
Denoising a 1024 × 1024 image pixel by pixel means handling over three million numbers at every step. That is slow and expensive.
Latent diffusion, described in High-Resolution Image Synthesis with Latent Diffusion Models, solved this and is the basis of Stable Diffusion and many successors.
- A separate network, an autoencoder, learns to compress images into a much smaller grid of numbers (the latent) and to decode that grid back into an image.
- Diffusion runs entirely on the small latent.
- The decoder turns the finished latent into the full-resolution picture.
Working on a representation dozens of times smaller is what brought image generation from data centres to consumer graphics cards.
The full pipeline
| Stage | What happens |
|---|---|
| 1. Encode the prompt | A text encoder turns your words into vectors |
| 2. Start from noise | A random latent is created from the seed |
| 3. Denoise in steps | The network repeatedly predicts and removes noise, guided by the prompt |
| 4. Decode | The autoencoder's decoder converts the latent into pixels |
| 5. Optional extras | Upscaling, face fixes, safety filters |
The denoising network itself was originally a U-Net, a convolutional design. Many newer models use transformers instead; see transformers explained.
Beyond text to image
The same machinery supports other tasks:
- Image to image. Start from an existing picture with some noise added, instead of pure noise, and denoise with a new prompt.
- Inpainting. Regenerate only a masked area.
- Outpainting. Extend a picture past its edges.
- Structural control. Guide generation with a sketch, a pose or a depth map.
- Personalisation. Small add-on weights (such as LoRA adapters) teach a model a particular style or subject.
- Video. Denoise a sequence of frames together so they stay consistent.
Why hands, text and counting go wrong
- Hands. They are small in most photos, appear in countless poses, and are often partly hidden. The model has learned "hand-like" texture better than hand anatomy. Recent models are much improved.
- Writing. Letters demand exact shapes in an exact order, and older text encoders worked with word pieces, not spelling. Newer models with stronger encoders render text far better.
- Counting and positions. "Three cats to the left of two dogs" requires precise binding of numbers and relationships, which these models handle loosely.
- Attribute mix-ups. "A red cube and a blue ball" can come out as a blue cube.
- Consistency. Each image is generated independently, so keeping a character identical across pictures takes extra techniques.
The model has learned statistical regularities of images. It has no model of bones, physics or grammar.
Other approaches
- GANs (generative adversarial networks) dominated before diffusion. A generator and a discriminator compete. They are fast but harder to train and less varied.
- Autoregressive models treat an image as a sequence of tokens and predict them one at a time, as language models do with text. Some recent multimodal systems generate images this way, or by combining both ideas.
Training data and wider questions
These models are trained on very large collections of images with captions, much of it gathered from the web. That has raised real and unresolved questions: copyright and the consent of artists, biases inherited from the data, and the use of generated images to deceive. Responses include licensing agreements, opt-out mechanisms, content filters, and provenance standards and watermarks that label AI-generated media. Training itself needs substantial computing resources; see why training AI needs so many GPUs.
Frequently asked questions
What is a diffusion model?
A neural network trained to remove noise from images. Run repeatedly from pure noise, it produces a new image.
Do AI image generators copy existing pictures?
They do not store a library of images and paste from it. They learn patterns from training data and generate new images, though heavily repeated training images can sometimes be reproduced closely.
What is a seed?
A number that determines the random starting noise. The same seed, prompt and settings give the same image.
What does the guidance scale do?
It controls how strongly the image is pushed to match the prompt. Higher values follow the prompt more literally at some cost in naturalness.
Conclusion
A diffusion model turns the simple skill of removing noise into the ability to create. Start from static, denoise a little at a time, and let the text prompt steer each step. Doing that in a compressed latent space made it fast enough for everyday use, and the same idea now drives image editing and video generation too.
Related articles
- How Embeddings Turn Words Into Numbers
- What Is Attention and Why Was It Such a Breakthrough?
- How Neural Networks Actually Learn
- Why Training AI Needs So Many GPUs
