About AI Image Video
Write or paste your script and every sentence becomes a scene with its own generated picture. The pictures are drawn to illustrate your actual words, animated with a slow zoom so they move like film rather than sitting there like a slideshow, and cut together in time with the narration.
Nothing here is stock. Every frame is generated, which means you are never limited to what a photographer happened to shoot β but it also means the imagery is synthetic, and you should say so wherever that matters. If you want real photographs instead, use Script to Video.
What you are paying for. Three things, and they are itemised above the button before you commit. The render is charged by the length of the finished video, because every render occupies a server for minutes. The pictures are charged one per scene, at 490 tokens each β the same per-image price as the AI Image Generator, from the same settings, so the two tools can never quietly disagree about what a picture is worth. A voice-over is optional and priced separately, per character of script.
The quote is a ceiling, never a floor. If the model declines to draw one of your lines, that scene falls back to a clean background, the video still finishes, and you are charged less than you were quoted. You are never charged more.
The pictures go into the video, not into your files. Only the finished MP4 is saved to your library. Twenty full-size images per render would quietly consume the storage you bought, so if you want images you can keep and reuse, generate them in the AI Image Generator instead β those are yours permanently.
How to use AI Image Video
Paste your script
Write it as normal prose. Each sentence becomes one scene and one picture, so punctuation is what controls both the pacing and the number of images you pay for.
Pick a picture style
Photographic, cinematic, illustration, watercolour or 3D. The style is applied to every scene, so the video holds together instead of changing look halfway through.
Add a voice-over (optional)
Choose a voice and the script is read aloud, with each picture held exactly as long as its line takes to say. Priced separately, per character of script.
Render and download
Watch the progress, then play it in the page or download the MP4.
Key Features
A picture drawn for every line
Each scene is generated from that scene's own sentence, so the imagery follows the script instead of approximating it with whatever stock footage happened to match a keyword.
One style, held throughout
The look you choose is applied to every scene in the video, which is what stops a set of individually good images from looking like five different videos spliced together.
You pay only for pictures that arrive
The price is quoted per scene up front and is a ceiling. Any scene the model declines to draw is refunded, and that scene falls back to a clean background rather than failing the whole render.
Captions burned in
Subtitles are drawn into the video itself, so they show on every player and on feeds that autoplay muted β which is most of them.
Common Use Cases
- Faceless YouTube and TikTok content where no stock footage fits the subject
- Explainers about abstract ideas that cannot be photographed
- Story, fiction and history narration that needs illustration rather than documentation
- Concept and pitch videos before anything real exists to film
- Consistent-looking series where every episode needs the same visual style
Security & Privacy
- Your script is sent over TLS to our image provider to draw the scenes, and to our speech provider if you ask for narration. It is not used to train anything.
- Scripts are screened before anything is charged. A script that cannot be turned into pictures is refused up front rather than after your tokens are spent.
- Finished videos are saved privately to your account and are never auto-deleted β only you or an admin can remove one.
- A finished video is served only to the account that made it. The link is not a credential anyone else can use.