Glossary
Diffusion model
A class of AI generative models that learn to reverse a noising process — starting from pure noise and progressively denoising into an image or video matching the prompt.
In depth
What it means in practice
Diffusion models are the dominant architecture for AI image generation as of 2026. Models including Stable Diffusion (SDXL), FLUX, and Nano Banana 2 are all diffusion-based. Video models like Sora 2 and Kling extend the diffusion idea to spacetime — denoising a noise volume to produce a coherent sequence of frames.
The core advantage of diffusion models over earlier architectures (GANs, VAEs) is sample quality and prompt fidelity. The trade-off is generation speed: each output requires multiple denoising iterations rather than a single forward pass. This is why diffusion-based generations cost what they do.
Related
Inference steps
The number of iterative refinement passes the model takes to produce an output. More steps generally means higher quality, with diminishing returns above a model-specific threshold.
CFG scale
A parameter controlling how closely the model follows the text prompt. Higher CFG = more literal interpretation; lower CFG = more creative freedom. Typical range 4-12.
Text-to-image
A generation mode where the input is text (the prompt) and the output is an image. The most common AI image generation mode.