Image generation model
Also called AI images · AI-generated images · text-to-image model
An AI model that turns a text prompt into a still image, the model layer behind AI-generated visuals in a production stack.
In more detail
An image generation model takes a written prompt, sometimes with a reference image or style setting, and returns one or more still images. It is billed per image (see cost per image) rather than per token, and it sits in a stack as its own replaceable layer, separate from the model writing the script and the tool assembling the final video.
Example
A documentary can mix AI-generated stills for scenes with no available real footage alongside licensed stock footage for scenes that do, rather than committing to one visual source for the entire video. That mix is a deliberate production choice, not a limitation of either source on its own.
Why it matters
Which images end up in a video, and how many of them are generated versus licensed, is one of the clearer signals platforms look at when assessing whether a channel's content is original and authentic rather than a reused or inauthentic content pattern. A stack that leans entirely on generated visuals with no other input reads differently than one that treats generation as one layer among several.