Glossary
In-image text
Also known as: text rendering, in-frame text
Text that appears inside the generated image as part of the scene — signage, packaging copy, quote cards, infographic labels. Historically a weakness of diffusion models; in 2026, only the strongest models render it legibly.
In depth
What it means in practice
For most of the 2022-2024 era of AI image generation, in-image text was approximately gibberish — text-shaped marks that looked like letters from a parallel alphabet rather than readable words. Nano Banana 2 (2025) and the latest FLUX checkpoints solved this for Latin scripts; Nano Banana 2 extends to non-Latin scripts (Chinese, Japanese, Arabic, Cyrillic) at class-leading fidelity.
For any output that needs legible text — quote cards, packaging mockups, signage in establishing shots, infographic data labels, social posts with text overlays — Nano Banana 2 is the only viable choice in the consumer bracket. Other models will produce the visual but lose legibility on the text.
Related
Identity preservation
A model's ability to keep a person's face or a character's appearance consistent across multiple generations from the same reference. Critical for multi-panel collages and character-driven content.
Output resolution
The pixel dimensions of a generated image or video. Common values: 1024×1024 (default for most image models), 4K (Nano Banana 2 portrait), 1080p (most video models).
Prompt
The text instruction given to an AI model that describes what to generate. For image and video models, prompts mix subject description, style direction, camera or composition hints, and optional negative prompts.