Generation and render
This page covers the two ways media gets made: generating still images and video clips from a prompt, and rendering a finished creative from a template.
Generation: images and video from a prompt
Generation produces a still image with Gemini 3.1 Flash Lite Image or a video clip with one of four video models from a text prompt, runs it on your Organization's own Google Cloud project, and drops the finished file into your asset library as a frozen, read-only asset.
It is built to run asynchronously: a request is queued, a worker carries it out, and video, which is slow, is polled to completion. Each job moves through queued, running, and then succeeded, failed, or canceled. A queued job can be canceled before anything is billed. A running job can be stopped too: the provider's work so far is abandoned (not unwound), and the job settles as canceled so a stuck generation stops blocking its scene, storyboard video, or render. A stop button appears wherever a generation or render is in flight.
A job or render with no progress for over five minutes counts as stuck: healthy work reports in more often than that, so silence means its background task chain was lost. When that happens a "Recover stuck work" action appears (on the campaign's Generate and Render matrix steps, and in the Render Queue): one click re-arms every lost poll and wait in your Organization, so work that actually finished settles with its real result and work that truly hung fails honestly with a retry offered. Recovery never bills anything new; only an explicit retry or re-render does.
What works today
The generation pipeline works end to end against live providers: image generation through Gemini 3.1 Flash Lite Image, and video generation through Gemini Omni or Veo, both billed to the Organization's Google Cloud project, with the output frozen into the asset library. Platform admins can watch generation jobs on the monitoring console, and jobs show up in the activity feed.
Generate from the Generate step
You trigger generation from a campaign's Generate step: the storyboard assembles into one video per clip of its clip plan, which is one clip unless you have split the scenes up. On the Storyboard step each scene can also get a fast still as an image preview; previews are only for you and never end up in a render. A separate, faster image model draws them, so a preview frame never matches the finished footage. Generation conditions on the sample items' linked images as references. The finished combination packages then go through the render step below.
The prompt the video model receives is composed from the storyboard: the direction block leads, each scene becomes a timed shot, and continuity rules, reference-image roles, and the campaign's must-haves follow. The Show the video prompt button in each clip's header on the Storyboard step opens exactly that text for that clip and the sample combination, so you can check what the model will be told before spending a generation. Every other combination gets the same prompt with its own item values and images filled in. The dialog also flags any {{variables}} that do not resolve, which would stop a generation from starting.
Choosing the video model
Every campaign picks the model its footage is generated with. The choice sits in the Production section of the Activation plan step, beside the target and the platform, and the Plan summary on the later steps echoes it. Campaign settings shows the same choice as a read-only echo; the voice reading the voiceover is still set there.
Four models are on offer, as cards you compare and pick from. None of them is a fallback for another; they trade length, cuts, resolution, and price.
| Gemini Omni 1.1 Flash | Gemini Omni Flash | Veo 3.1 | Veo 3.1 Fast | |
|---|---|---|---|---|
| Clip length | any whole second from 3 to 40 | 3 to 10 seconds | 4, 6, or 8 seconds, extendable to 29 | Same as Veo 3.1 |
| Scenes in one clip | cuts between scenes by itself | cuts between scenes by itself | one scene per clip reads best | one scene per clip reads best |
| Resolution | 720p, 1080p, or 4K | 720p | 720p, 1080p, or 4K | 720p, 1080p, or 4K |
| Reference images | up to 7 | up to 7 | up to 3, which lock the clip to an 8 second base | up to 3, same lock |
| Price per second | $0.10 at 720p, $0.15 at 1080p, $0.30 at 4K | $0.10 | $0.20 at 720p and 1080p, $0.40 at 4K | $0.08 at 720p, $0.10 at 1080p, $0.25 at 4K |
Gemini Omni 1.1 Flash is the platform default. It cuts between scenes inside a single generation, so a three-scene story can be one clip, and it reaches 40 seconds through the extend chains described in The clip plan. Gemini Omni Flash is the older model of the same family, kept so campaigns that started on it can finish on it; it stops at 720p and 10 seconds. Veo 3.1 is the quality tier, and Veo 3.1 Fast is the same shape for less money.
A change of model applies to the next thing you generate. Clips that already exist stay as they are, and nothing re-renders on its own. If the new model does not produce a length the clip plan already uses, those clips are marked as not generatable with the length named, and you pick a new one on the storyboard. Klyo does not substitute a length silently.
Resolution
The resolution row appears under the cards once the chosen model lets you set one. It offers the tiers that model serves at your campaign's orientation, which comes from the bound template, plus the platform default. Gemini Omni Flash serves 720p only and ignores a resolution request, so the row is hidden for it entirely rather than offering a knob that does nothing.
Gemini Omni 1.1 Flash, Veo 3.1, and Veo 3.1 Fast all serve 720p, 1080p, and 4K at both landscape and portrait, verified live on the endpoint production uses. 4K costs about three times 720p, so pick it for a hero cut, not for a whole matrix.
What a second costs
The prices in the table are the provider's own list prices for video without audio, which is what Klyo generates: the voiceover is a separate track mixed in at render. They are not what your organization is charged, and they move when Google moves them.
Seconds and the resolution tier are what you pay for, so the arithmetic is simple. A 22 second Veo 3.1 clip at 720p is about four times a 10 second Gemini Omni 1.1 Flash clip at 720p, and it costs that once per combination in the matrix, not once per campaign. The confirm dialog before a fan-out does the multiplication for you.
Voiceover
When a template's audio field binds the storyboard voiceover, Klyo generates a narrator track from the storyboard itself. Each scene's voiceover line is the script: Gemini text-to-speech reads them, in scene order, into one spoken track.
You set the binding on the Package step, where an audio field's fill source offers "Storyboard voiceover" alongside pinning a library track. The voiceover speaks the campaign's language, set in Campaign settings alongside the voice. A language change applies on the next generation, not to a track already made, so regenerate to hear it.
On the Generate step a voiceover card appears once the binding is set. Generate the sample voiceover there, play it back, and regenerate for another take (a regenerate bills the text-to-speech again). The sample render waits on this voiceover, so generate it before rendering the sample.
On the matrix fan-out every combination generates its own voiceover from its own resolved scene lines, so a script with {{variables}} reads each cell's real values. Generation refuses a line that still carries an unresolved variable rather than speaking the literal braces, so fix or remove any leftover token in the storyboard first.
Render: a finished creative from a template
Rendering takes a combination package, its template version, and its filled inputs, and produces a finished video or image through Creatomate.
Rendering is built
Rendering is connected (record 0081) and fans out per combination (record 0099). From a campaign's Render matrix step you trigger the sample proof render that checks the template, then fan the matrix out. A template whose video field binds the storyboard generates its own fresh clip per combination, one per clip of the plan, composed from that cell's own catalogue items; every cell also carries its own catalogue item images and ad copy. The finished video freezes into the asset library. The Review console plays each Final Video so you can approve, reject, or re-render it, one cell or in bulk. Approved combinations then activate: pushed to YouTube or handed off as a Google Ads or DV360 pack (Activation). The cross-channel roll-up over the pulled metrics is the frontier still ahead.
The cost estimate
Before you fan the matrix out, the confirm dialog estimates what the run will spend. It is a rough figure, flat per-item rates for renders and voiceovers and the chosen model's per-second rate for video, not live provider pricing, broken into up to three lines:
- Renders: one render credit per combination, always present.
- Generated clips: the plan's clips per combination, shown when a template video field binds the storyboard, priced by the chosen model and resolution.
- Generated voiceovers: one narrator track per combination, shown when an audio field binds the storyboard voiceover.
The sample is already rendered, so the estimate covers the rest of the matrix. The real bill depends on each clip's length and the provider's current rates. The spend lands on your organization's own provider accounts (its Google Cloud project for inference and its own render subscription), never on a personal card.
How the pieces fit
Storyboard ──generate──▶ a fresh video per combination
+ per-scene image previews (never in a render)
Combination Package + Template version + inputs ──render──▶ Final Video (Creatomate)
Final Video ──review──▶ approved / rejected / re-renderedThe design principle behind both is "own the engine, rent the render": Klyo keeps the orchestration, prompts, and assets, and rents only the raw generation and render steps from providers it can swap out later.
What you can do today
- Generate a scene's image preview, or every scene's at once, from the Storyboard step, conditioned on the sample items' linked images as references.
- Run image and video generation against live providers (through the API and the admin monitoring console), with results frozen into the asset library.
- Choose a campaign's video model (Gemini Omni 1.1 Flash, Gemini Omni Flash, Veo 3.1, or Veo 3.1 Fast) and the resolution it generates at in the Production section of the Activation plan step (Gemini Omni Flash fixes 720p, so it offers no resolution row), and the voice reading the voiceover in Campaign settings.
- Generate a narrator voiceover from the storyboard's scene lines when an audio field binds it, one track for the sample and a fresh one per combination, in the campaign's language set in Campaign settings.
- Render the sample and fan the matrix out into Final Videos through Creatomate, getting a freshly generated clip per combination when the template's video field binds the storyboard.
- Approve, reject, or re-render each Final Video, one cell or in bulk, from the Render Queue or the Review console.
- Inspect template structure and history (see Templates).
Not available yet
- The cross-channel roll-up that joins the pulled metrics across networks into one comparison. The nightly KPI pull that feeds it is built for owned YouTube uploads and for Tracked Campaigns and ships switched off by default; the roll-up over that data is the frontier still ahead.

