Gemini Omni 1.1 Flash Bills Video by the Second and Forgets Your Product Every Ten of Them
📢
← Back to Blog

Gemini Omni 1.1 Flash Bills Video by the Second and Forgets Your Product Every Ten of Them

John Aspinall · · 8 min read

For the first time, you can price an Amazon video before you make it. Google's video model now bills by the second at a rate that depends on the resolution you export, and Amazon's video surfaces have fixed specs. Put those two facts next to each other and a 30-second Sponsored Brands Video at 1080p is $4.50 of generation. A 40-second one is $6. Six 10-second drafts to pick a direction from is $1.80.

That is not the interesting part. I told you in May that the cost of listing video had collapsed, and it had. The interesting part of this release is the control layer that arrived with the price list, and the one line in the documentation that decides whether any of it is usable on a detail page: when the model extends a clip, it remembers the previous ten seconds and nothing before that.

What happened

On August 27, 2026 Google moved Gemini Omni 1.1 Flash to general availability in the Gemini API with video extension, first-and-last-frame interpolation and resolution control, per the Gemini API changelog and the updated model card. Individual outputs run 3 to 10 seconds; you extend in 10-second increments to a 40-second cumulative cap; you can hand it two still images and have it interpolate a continuous clip between them; there is a 360p drafting mode; and per-second output pricing is reported at roughly $0.03 (360p), $0.10 (720p), $0.15 (1080p) and $0.30 (4K), with the 1080p and 4K tiers being upscales of a 720p render rather than native generation (XenoSpectrum's breakdown, and Google's own pricing page documents the 720p token rate that nets to about ten cents a second).

Hold the per-second figures loosely. Google publishes token rates, not a clean dollar-per-second table, and the conversions above are derived. The shape is not in dispute: you pay per second, you pay more per second the larger you export, and the two largest tiers are resizes.

Why most brand owners will read this wrong

The dumb take is "video is free now, put video on everything." Video was already close to free in May. Nothing about your listing improved because generation got cheaper, and a bad 40-second clip on a detail page costs you exactly what a bad 40-second clip cost you before, which is the shopper's attention and possibly the sale.

The second dumb take is the one with a price attached: "finally, 4K product video." Amazon's Sponsored Brands Video spec accepts 1280×720, 1920×1080 and 3840×2160 and recommends 1080p, per Amazon Ads' own spec page. The model renders at 720p internally. Paying $0.30 a second for 4K buys you a 720p image stretched twice, wrapped in a file that says 4K, and the first person to notice is the shopper who taps to full screen on a tablet and sees soft edges on your label.

The real signal is smaller and more useful. This is the first video model whose documentation reads like a spec sheet you can lay directly against Amazon's spec sheet. Segment length, cumulative length, aspect ratio, the resolution ladder, and a stated memory window. That means the questions that used to be "can it do this" are now "what does this cost and where does it break," and both of those have answers.

What actually changes for someone running $200K/mo on Amazon

You can price a variant before you brief it. A 20-second SBV at 1080p is $3 of compute. Testing three hooks against each other is under $10. That number is so small it is no longer a budgeting input, which moves the entire constraint to the brief and the QA, where it always belonged. The cost in this business is the person who checks the output, and that person did not get cheaper on Wednesday.

The 40-second clip is four 10-second clips wearing a trench coat. Extension is at the end of a clip only, no mid-scene insertion, no backward extension, and the model references only the preceding ten seconds when it continues. For a lifestyle reel that is fine. For a listing video the product at second 35 has to be the product at second 2, and the mechanism that would guarantee that does not exist. Drift compounds across segments: a slightly different cap, a label that migrated, a colourway that warmed up. The model card says it plainly: maintaining consistency through edits and rendering accurate text remain a challenge.

Keyframes are the fix, and they're the actual feature. First-and-last-frame interpolation means you can anchor every segment boundary with a real photograph of your real product. Hero on white as the first frame, your in-use shot as the last, then the last frame of segment one becomes the first frame of segment two. The product is pinned at second 0, 10, 20, 30 and 40 by images you own. What the model invents is the motion between pins, which is the part it is good at.

Muted autoplay makes the audio irrelevant on the surface you'd use first. SBV autoplays on mute with the toggle in the lower right. Omni's audio continuation is a nice feature for a lifestyle reel and does nothing for a search-results placement. Brief the clip to work silent and caption it.

The reference-video limit is a real ceiling. Up to three reference videos, no cross-referencing between them. You cannot hand it three angles of your product and expect it to reconcile them into one object. One reference, one product, one segment at a time.

SynthID rides along. Every output is watermarked. That is a provenance column, not a compliance problem, and it belongs in the same automation inventory as your text models: tool, model string, pinned or floating, current rate, re-check date, retirement date, watermark.

What I'd do this week if I were them

  1. Keyframe from your own photography at every ten-second boundary. Never let the model carry the product across a segment. Your hero, your slot-two frame, your in-use shot, in that order, as pins. If you do only one thing from this post, do this.
  2. Draft at 360p, judge at 720p, deliver at 1080p, never buy 4K. Google's recommended workflow is the same shape. A 10-second 360p draft is thirty cents; run six, kill five, then spend the ten cents a second on the survivor.
  3. Write the brief as four 10-second sentences. One shot per sentence, one thing happening per shot. The model remembers one segment back, so a brief that expects it to remember the opening in segment four is a brief that will produce drift and blame the tool.
  4. Build a six-clip golden set from six real SKUs and check the label at every boundary. Freeze-frame at 0, 10, 20, 30, 40 and read the text on the product. Any drift is a fail. This is a twenty-minute QA and it is the difference between a listing video and a compliance claim on a live detail page.
  5. Put it on Sponsored Brands Video first, not the PDP. SBV gives you a CTR read in two to three weeks and the traffic is muted, short and skimmed, which is the most honest test of a clip that has to work silently. PDP video moves CVR on a 30-to-60-day window, and eight weeks from peak that window closes in mid-September. Ship the SBV test now or freeze it until January.

What I'd ignore

The 40-second headline. It is a cap on concatenation, not on memory, and memory is the number that matters for a product.

The 4K headline. It's a resize with a price tag.

The Omni-versus-Veo-versus-Kling ranking cycle. No published evaluation measures whether a generated clip keeps your label legible at second 27. That benchmark is still the one you build yourself from your own SKUs.

The argument about whether this replaces Amazon's own Creative Studio video generator. Amazon's tool is free, uses your product images and lives in the console where the campaign lives. For a brand that has never run SBV, that is still the right first stop. Omni 1.1 earns its place when you need a specific shot the template tool won't give you and you're willing to pin every segment yourself.

The Gemini-app and Flow consumer coverage. Different product, different limits, no keyframing workflow you'd trust with a catalog.

In May the story was that the price of listing video went away. This week the story is that the controls arrived, and they arrived with a memory window written on the box. The brands that get real value out of this will be the ones who read that line and built the QA around it. The ones who read the 4K line will pay double for a blurrier file and wonder why the reviews mention it.

Install this as an agent, not a checklist.

The Operator Intelligence: Multi-Agent OS cohort is a 4-week live build: 2-3 specialist agents with their own seats, running real workflows on your actual catalog. Starts Mon, Sep 14 · $499 · 12 seats · replays included.

See the cohort →

Want to see it working first? Watch the free replay — the whole system built live on a real ecommerce business.