Skip to content

Generation and render ​

This page covers the two ways media gets made: generating still images and video clips from a prompt, and rendering a finished creative from a template.

Generation: images and video from a prompt ​

Generation produces a still image with Gemini 3.1 Flash Lite Image or a video clip with one of four video models from a text prompt, runs it on your Organization's own Google Cloud project, and drops the finished file into your asset library as a frozen, read-only asset.

It is built to run asynchronously: a request is queued, a worker carries it out, and video, which is slow, is polled to completion. Each job moves through queued, running, and then succeeded, failed, or cancelled. A queued job can be cancelled before anything is billed. A running job can be stopped too: the provider's work so far is abandoned (not unwound), and the job settles as cancelled so a stuck generation stops blocking its scene, storyboard video, or render. A stop button appears wherever a generation or render is in flight.

A job or render with no progress for over five minutes counts as stuck: healthy work reports in more often than that, so silence means its background task chain was lost. When that happens a "Recover stuck work" action appears (on the campaign's Generate and Render matrix steps, and in the Render queue): one click re-arms every lost poll and wait in your Organization, so work that actually finished settles with its real result and work that truly hung fails honestly with a retry offered. Recovery never bills anything new; only an explicit retry or re-render does. The Generate step's banner counts every kind of work that step owns, storyboard clips and the sample render alike, so one click covers all of it.

Work that never reports back at all is not left running: a render whose polling window runs out, a video chain stranded between two segments, and a Creative Score whose attempts are spent are each settled by a background sweep, with the reason recorded, so nothing sits in a running state for good.

What works today

The generation pipeline works end to end against live providers: image generation through Gemini 3.1 Flash Lite Image, and video generation through Gemini Omni or Veo, both billed to the Organization's Google Cloud project, with the output frozen into the asset library. Platform admins can watch generation jobs on the monitoring console, and jobs show up in the activity feed.

Generate from the Generate step

You trigger generation from a campaign's Generate step: the storyboard assembles into one video per clip of its clip plan, which is one clip unless you have split the scenes up. On the Storyboard step each scene can also get a fast still as an image preview; previews are only for you and never end up in a render. A separate, faster image model draws them, so a preview frame never matches the finished footage. Video generation is conditioned on the reference images each clip ranks for itself, also on the Storyboard step. The finished combination packages then go through the render step below.

The prompt the video model receives is composed from the storyboard: the direction block leads, each scene becomes a timed shot, and continuity rules, reference-image roles, and the campaign's must-haves follow. The Show the video prompt button in each clip's header on the Storyboard step opens exactly that text for that clip and the sample combination, so you can check what the model will be told before spending a generation. Every other combination gets the same prompt with its own item values and images filled in. The dialog also flags any {{variables}} that do not resolve, which would stop a generation from starting.

Choosing the video model ​

Every campaign picks the model its footage is generated with. The choice sits in the Production section of the Activation plan step, beside the target and the platform, and the Plan summary on the later steps echoes it. Campaign settings shows the same choice as a read-only echo; the voice reading the voiceover is still set there.

The models on offer sit in one table, a row each, with the facts a choice turns on in their own columns: whether the model cuts between scenes itself, the clip lengths it reaches, the highest resolution it serves, and the provider's list price a second, plus how many reference images it reads. Sort by any column (cheapest first, longest first), search by name, and, once your organization has Partner models too, filter or group by provider. Click a row to pick it.

Once the campaign has a storyboard, a Fits column reads every model against the clip plan as it stands. No model is ever refused: the column says what changes if you pick it. A clip whose length the model does not produce generates at the nearest length it does, and the render trims the extra back to the placement; a clip holding more scenes than the model cuts cleanly generates as shots described in order, with the model choosing where the cuts fall; a clip staging more reference images than it reads drops the overflow. The one reading worth acting on is a clip longer than anything the model reaches, which leaves a gap in the render. Hover for which clip and why.

None of the models is a fallback for another; they trade length, cuts, resolution, and price. The four Google models:

Gemini Omni 1.1 FlashGemini Omni FlashVeo 3.1Veo 3.1 Fast
Clip lengthany whole second from 3 to 403 to 10 seconds4, 6, or 8 seconds, extendable to 29Same as Veo 3.1
Scenes in one clipcuts between scenes by itselfcuts between scenes by itselfone scene per clip reads bestone scene per clip reads best
Resolution720p, 1080p, or 4K, plus 360p in Draft mode720p720p, 1080p, or 4K720p, 1080p, or 4K
Reference imagesup to 7up to 7up to 3, which lock the clip to an 8 second baseup to 3, same lock
Price per second$0.034 at 360p, $0.10 at 720p, $0.15 at 1080p, $0.30 at 4K$0.10$0.20 at 720p and 1080p, $0.40 at 4K$0.08 at 720p, $0.10 at 1080p, $0.25 at 4K

Gemini Omni 1.1 Flash is the platform default. It cuts between scenes inside a single generation, so a three-scene story can be one clip, and it reaches 40 seconds through the extend chains described in The clip plan. Gemini Omni Flash is the older model of the same family, kept so campaigns that started on it can finish on it; it stops at 720p and 10 seconds. Veo 3.1 is the quality tier, and Veo 3.1 Fast is the same shape for less money.

The reference-image row is a hard limit, and on Veo it is small enough to matter: three images, which a two-axis matrix can fill with catalogue item images alone. Each clip's reference priority decides which images fill those slots; anything past the number is staged and never sent.

A change of model applies to the next thing you generate. Clips that already exist stay as they are, and nothing re-renders on its own. If the new model does not produce a length the clip plan already uses, those clips are marked as not generatable with the length named, and you pick a new one on the storyboard. Klyo does not substitute a length silently.

Resolution ​

The resolution row appears under the table once the chosen model lets you set one. It offers the tiers that model serves at your campaign's orientation, which comes from the bound template, plus the platform default. Gemini Omni Flash serves 720p only and ignores a resolution request, so the row is hidden for it entirely rather than offering a knob that does nothing.

Gemini Omni 1.1 Flash, Veo 3.1, and Veo 3.1 Fast all serve 720p, 1080p, and 4K at both landscape and portrait, verified live on the endpoint production uses. 4K costs about three times 720p, so pick it for a hero cut, not for a whole matrix.

Draft mode ​

Draft mode generates every clip at 360p instead of the resolution above, for about a third of the cost of 720p. Use it while you are still changing the storyboard: you get a watchable clip that shows whether the scenes read, the pacing works, and the tokens resolve, without paying the finish price for an answer you already know you will regenerate.

The switch sits at the top of the Generate step, in the "This run uses" card beside the gear, and the Resolution line reads "Draft (360p)" while it is on. It is a mode, not a resolution choice: turn it off and the resolution you picked in Production comes back exactly as it was. Only Gemini Omni 1.1 Flash serves the draft tier, so the switch is not offered on the other models.

Nothing already generated changes when you flip it. Draft mode applies to the next run, so regenerate the clip to see it at the new tier.

What a second costs ​

The prices in the table are the provider's own list prices for video without audio, which is what Klyo generates: the voiceover is a separate track mixed in at render. They are not what your organization is charged, and they move when Google moves them.

Seconds and the resolution tier are what you pay for, so the arithmetic is simple. A 22 second Veo 3.1 clip at 720p is about four times a 10 second Gemini Omni 1.1 Flash clip at 720p, and it costs that once per combination in the matrix, not once per campaign. The confirm dialog before a fan-out does the multiplication for you.

Voiceover ​

When a template's audio field binds the storyboard voiceover, Klyo generates a narrator track from the storyboard itself. Each scene's voiceover line is the script: Gemini text-to-speech reads them, in scene order, into one spoken track.

You set the binding on the Package step, where an audio field's fill source offers "Storyboard voiceover" alongside pinning a library track. The voiceover speaks the campaign's language, set in Campaign settings alongside the voice. A language change applies on the next generation, not to a track already made, so regenerate to hear it.

On the Generate step a voiceover card appears once the binding is set. Generate the sample voiceover there, play it back, and regenerate for another take (a regenerate bills the text-to-speech again). The sample render waits on this voiceover, so generate it before rendering the sample.

On the matrix fan-out every combination generates its own voiceover from its own resolved scene lines, so a script with {{variables}} reads each cell's real values. Generation refuses a line that still carries an unresolved variable rather than speaking the literal braces, so fix or remove any leftover token in the storyboard first.

Render: a finished creative from a template ​

Rendering takes a combination package, its template version, and its filled inputs, and produces a finished video or image through Creatomate.

Rendering is built

Rendering is connected (record 0081) and fans out per combination (record 0099). From a campaign's Render matrix step you trigger the sample proof render that checks the template, then fan the matrix out. A template whose video field binds the storyboard generates its own fresh clip per combination, one per clip of the plan, composed from that cell's own catalogue items; every cell also carries its own catalogue item images and ad copy. The finished video freezes into the asset library. The sample proof belongs to the template it rendered, so switching the campaign's template asks for a fresh sample render before the matrix opens again. It also reads Out of date once the sample combination, a field binding, or the storyboard has moved since it rendered: hover the chip to see which. That is a warning, never a block, and nothing re-renders itself, because a re-render spends. The same idea runs backwards through the builder: the Ideas, Storyboard and Ad copy steps each show a chip when the thing they were generated from has changed since, with the one action that clears it.

A combination the fan-out cannot start (an axis item with no image, say) is written as a failed cell with its reason and a Retry, and counts on the step's needs-attention total, rather than disappearing into a failure count. The Review console plays each Final Video so you can approve, reject, or re-render it, one cell or in bulk. Approved combinations then activate: pushed to YouTube or handed off as a Google Ads or DV360 pack (Activation). The cross-channel roll-up over the pulled metrics is the frontier still ahead.

The cost estimate ​

Before you fan the matrix out, the confirm dialog estimates what the run will spend. It is a rough figure, flat per-item rates for renders and voiceovers and the chosen model's per-second rate for video, not live provider pricing, broken into up to three lines:

  • Renders: one render credit per combination, always present.
  • Generated clips: the plan's clips per combination, shown when a template video field binds the storyboard, priced by the chosen model and resolution.
  • Generated voiceovers: one narrator track per combination, shown when an audio field binds the storyboard voiceover.

The sample is already rendered, so the estimate covers the rest of the matrix. The real bill depends on each clip's length and the provider's current rates. The spend lands on your organization's own provider accounts (its Google Cloud project for inference and its own render subscription), never on a personal card.

When your organization's credits have a hard stop and the run would need more credits than are left, the render is refused before anything starts. Ask a platform admin to add credits, or wait for the Monday reset.

How the pieces fit ​

Storyboard ──generate──▶ a fresh video per combination
                       + per-scene image previews (never in a render)
Combination Package + Template version + inputs ──render──▶ Final Video (Creatomate)
Final Video ──review──▶ approved / rejected / re-rendered

The design principle behind both is "own the engine, rent the render": Klyo keeps the orchestration, prompts, and assets, and rents only the raw generation and render steps from providers it can swap out later.

What you can do today ​

  • Generate a scene's image preview, or every scene's at once, from the Storyboard step, conditioned on the campaign-wide reference images you have ticked.
  • Rank the reference images each clip generates against, so the ones that matter reach the video model before it hits its limit.
  • Run image and video generation against live providers (through the API and the admin monitoring console), with results frozen into the asset library.
  • Choose a campaign's video model (Gemini Omni 1.1 Flash, Gemini Omni Flash, Veo 3.1, or Veo 3.1 Fast) and the resolution it generates at in the Production section of the Activation plan step (Gemini Omni Flash fixes 720p, so it offers no resolution row), and the voice reading the voiceover in Campaign settings.
  • Turn on Draft mode from the Generate step to iterate at 360p for about a third of the 720p cost, on Gemini Omni 1.1 Flash, and turn it off to get your chosen resolution back unchanged.
  • Generate a narrator voiceover from the storyboard's scene lines when an audio field binds it, one track for the sample and a fresh one per combination, in the campaign's language set in Campaign settings.
  • Render the sample and fan the matrix out into Final Videos through Creatomate, getting a freshly generated clip per combination when the template's video field binds the storyboard.
  • Approve, reject, or re-render each Final Video, one cell or in bulk, from the Render queue or the Review console.
  • Inspect template structure and history (see Templates).

Not available yet ​

  • The cross-channel roll-up that joins the pulled metrics across networks into one comparison. The nightly KPI pull that feeds it is built for owned YouTube uploads and for Tracked Campaigns and ships switched off by default; the roll-up over that data is the frontier still ahead.

A Digitl product. This guide covers what is built today, not the full roadmap.