Video generation model
Also called text-to-video model · AI video generator
An AI model that generates a video clip directly from a text prompt or a still image, as distinct from a still-image model or licensed stock footage.
In more detail
Where an image generation model outputs a single still, a video generation model outputs a short moving clip, typically a few seconds long, and is billed accordingly, usually per second or per clip rather than per image. It is a newer and generally more expensive layer than image generation, and most working faceless-channel stacks documented on this site still favour a mix of AI-generated stills and licensed stock footage over generated video, largely on cost and consistency grounds.
Example
A ten-second generated clip can cost many times more than a single generated still covering the same span of screen time in a slideshow-style edit, which is why swapping stills for generated video is a real cost decision and not a straightforward upgrade.
Why it matters
The gap between a still-image stack and a video-generation stack is currently more about cost and reliability than about which looks better, so the choice is worth treating as an economics question specific to a channel's volume rather than assuming the newer technology is automatically the right default.