Glossary
Text-to-image
Also known as: T2I, txt2img
A generation mode where the input is text (the prompt) and the output is an image. The most common AI image generation mode.
In depth
What it means in practice
Text-to-image is the default for most image models. The model takes a written description and produces an image matching it. Nano Banana 2 supports text-to-image on hilens, with HiDream-O1-image-1.5 announced to follow.
Text-to-image contrasts with image-to-image (input is an existing image plus optional prompt, model produces a modified version) and image-to-video (input is an image, model produces a video). The same model often supports multiple modes through different parameter settings rather than as distinct models.
Related
Prompt
The text instruction given to an AI model that describes what to generate. For image and video models, prompts mix subject description, style direction, camera or composition hints, and optional negative prompts.
Image-to-video
A generation mode where the input is a still image and the output is a video animating that image. Often paired with a text prompt describing the desired motion.
Reference image
An image uploaded alongside the prompt to condition the generation. The model preserves visual properties — composition, color palette, subject likeness, framing — from the reference in the output.