Most modern text-to-image systems are diffusion models. They generate images by gradually removing noise.
How Diffusion Works
During training, the model sees images with increasing amounts of random noise added and learns to predict and remove that noise. To generate, it starts from pure noise and applies many denoising steps, guided by a text prompt, until an image emerges.
Many systems work in a compressed latent space rather than on full-resolution pixels, which makes generation much faster.
How Text Guides the Image
A text encoder turns the prompt into an embedding that conditions each denoising step. A setting often called guidance scale controls how strongly the image follows the prompt versus looking natural.
What You Can Control
- Prompt and negative prompt (what to include and avoid).
- Seed for reproducibility.
- Number of steps (quality versus speed).
- Image-to-image and inpainting: edit or extend existing images.
- Conditioning tools that follow sketches, poses or depth maps.
Limitations
Text rendering, hands, counting objects and precise spatial layouts have historically been weak points, though they are improving.
Legal and Ethical Considerations
- Copyright and licensing questions around training data and outputs.
- Likeness and consent when depicting real people.
- Misinformation and deceptive imagery.
- Bias in how people and cultures are represented.
Check the model's licence and your organisation's policy, and label AI-generated images where appropriate.