Gemini Omni 1.1 Flash: What 40-Second 4K Video Means for Agency Content

Published August 28, 2026By ABD Legacy LLC
Gemini video API / AI video tools for agencies / best AI video model for ads

1. What changed on August 27

On August 27, 2026, Google made Gemini Omni 1.1 Flash generally available through the Gemini video API (model ID gemini-omni-1.1-flash in Google AI Studio), the production successor to the gemini-omni-flash-preview endpoint developers have used since June 30 [1][6]. The headline change for agencies: clips can now be extended in 10-second increments up to a cumulative 40 seconds, with the model using up to 10 seconds of prior video as context — so a 30-second ad or a looping brand asset can be directed as one scene instead of assembled from separate clips [1][2]. The rest covers the three capabilities that matter, Gemini Omni pricing per second, and where it fits in your AI video tools for agencies stack alongside our Gemini model pricing coverage.

2. The three capabilities that matter

1. 40-second scenes with a 10-second lookback. Previous Omni previews referenced only the final second when extending; Omni 1.1 Flash analyzes up to 10 seconds of prior context for visual consistency and narrative flow [2]. A 40 second AI video is now a realistic deliverable — built as 10s increments rather than a single long generation [1][2].

2. First/last frame control. Specify a starting and ending frame, and the model generates continuous video between the keyframes — the primitive behind camera orbits, zoom transitions, and seamless loops [1][2]. Prompt tags <FIRST_FRAME>/<LAST_FRAME> bind images to those roles [2].

3. A 360p draft → 4K upscale ladder. Output resolution runs 360p (draft) → 720p (default) → 1080p (upscaled) → 4K (upscaled) [2][5]. Important accuracy point: 1080p and 4K are upscales of the generated frames, not native renders — Google's docs label them as "upscaled" [2][4]. If you are evaluating a 4K AI video generator, "4K" means a 720p render upscaled [5]. Draft-tier economics (§3) make the workflow clear: iterate at 360p, render the keeper at 720p or upscaled 1080p/4K [1][5].

3. What Gemini Omni 1.1 Flash costs

Gemini Omni pricing is token-based, billed through the Gemini API: input at $1.50 per 1M tokens, text output at $9.00 per 1M, video output at $17.50 per 1M [3]. Video bills at 5,792 tokens per second of 720p — roughly $0.10 per second — so the AI video generation cost per second is the planning unit for client budgets, and a 40-second 720p scene runs about $4.00–$4.06 in output [3]. There is no free tier [3]. A 10-second 360p draft costs ~$0.30–$0.34 versus ~$1.01 at 720p, because 360p renders up to 60% faster at about one-third the cost [1][5].

TierResolutionPrice per 10s clipPrice per 40s sceneNotesSource
Draft360p~$0.30–$0.34~$1.20–$1.36up to ~60% faster than 720p, ~1/3 cost[1][5]
Standard720p~$1.01~$4.00–$4.06default; official per-second rate ~$0.10/s[3]
Upscaled1080p~$1.50 (reseller est.)~$6.00 (reseller est.)official rate not yet published by Google[5]
Upscaled4K~$3.00 (reseller est.)~$12.00 (reseller est.)official rate not yet published by Google[5]

1080p/4K per-second rates are marketplace/reseller estimates (~$0.15/s and ~$0.30/s respectively), not Google's published numbers [5]. Per-clip totals are arithmetic on the cited rates (analysis, not new sourced claims).

The 360p tier is where the economics get interesting: Apidog's worked example puts a 12-draft + 1-final-render workflow at $13.18 if every frame renders at 720p versus $5.07 if drafts run at 360p and only the keeper renders at 720p — about 62% less spend [5]. For context on what agencies charge for AI work overall, see how much an AI agency costs in 2026, and for video-as-a-service retainers, our AI agency pricing negotiation strategies.

4. The Gemini video API for agencies

For agencies building client pipelines, the API story is the practical part. Editing and generation run through the Interactions API — a conversational, multi-turn editing loop where each request references the previous previous_interaction_id, and response_format pins resolution and aspect ratio (9:16 or 16:9) [1][2]. Multimodal input (text, image, audio, video in one call), video references up to 3 clips × 3s each, subject/character reference images, text-in-video rendering, timecode prompting, and native synchronized audio on every clip are all supported [2].

Surfaces: the Gemini API and Gemini Enterprise Agent Platform (Adobe, Figma Weave, GMI Cloud, Runway); consumer-side, Google Flow carries the full feature set for AI Plus/Pro/Ultra subscribers, and scene extension is live in the Gemini app [1][4]. On API quality alone it's a genuine text to video API 2026 contender for short-form work — see our AI workflow automation implementation guide for slotting model APIs into client pipelines.

5. Ads and social video: the workflow change

This is where the update changes a recommendation. A 30-second ad — the most common paid-social creative length — now fits in one directed scene: a beginning frame, an ending frame, and the model fills the middle [1][2]. Combined with the 360p draft tier, the client iteration loop becomes: storyboard and test hooks at 360p (~$0.30–$0.34 per 10s draft), then render the keeper at 720p or upscaled 1080p/4K [1][5]. For ad testing that previously meant five-figure production budgets, a creative variant now costs under a dollar of API spend.

That economics shift is why, for best AI video model for ads decisions, Gemini Omni 1.1 Flash is now the default recommendation for volume and iteration — with the §7 caveat that Veo remains the fidelity pick. For the paid-social angle, our ChatGPT ads for AI agencies piece covers adjacent ad-creative workflows, and AI automation ROI for small business frames justifying the tooling cost to clients.

6. Where it still falls short

Vendor claims are strong; independent proof is not. No third-party benchmarks exist yet for Gemini Omni 1.1 Flash on any public video leaderboard — speed and quality figures are vendor-reported [4].

Operational limits that matter for client work [2]:

For agencies selling 4K AI video generator capability to clients, the upscaled-4K framing matters: an upscale, not native-4K rendering [2][5].

7. Veo 3.1 vs Gemini Omni: which to recommend

Google's docs position Gemini Omni Flash as the recommended default for video generation, with Veo as the higher-fidelity specialist — complementary tiers on the same API key, not rivals [2][5]. The practical split: Omni for volume, iteration, and conversational editing; Veo 3.1 for maximum fidelity or longer runtime (up to 148 seconds) [5].

CapabilityGemini Omni 1.1 FlashVeo 3.1
Price @720p~$0.10/s [3]~$0.40/s [5]
Price @4Kpending (reseller est. ~$0.30/s) [5]~$0.60/s [5]
Max scene length40s cumulative (10s increments) [1][2]up to 148s [5]
Native audioyes [2]yes [5]
Conversational editing (Interactions API)yes [1][2]not stated
Free tierno [3]not stated
Best fitads/social short-form, product demos, loops, client iterationmax fidelity + long runtime

Also on the same key: Veo 3.1 Fast at $0.10/s ($0.30/s @4K) and Veo 3.1 Lite at $0.05/s (no 4K) [5]. If the deliverable is 15–30 second social creatives, Omni's draft-to-final economics win; for long-form brand film, Veo 3.1's 148-second ceiling beats Omni's 40-second cumulative cap. Recommendation update: for ads/social video, Gemini Omni 1.1 Flash becomes the default for volume + iteration; keep Veo 3.1 where max fidelity and >40s runtime are required [2][5]. A full row-by-row breakdown lands on the AI video tools comparison page.

8. Bottom line for agencies

Switch now if you produce short-form social ads, product demos, and looping brand assets — the 40-second ceiling plus 360p drafting changes the cost math for iteration-heavy creative work [1][5]. Wait if you sell long-form brand film (Veo 3.1's 148s ceiling is still the fit), serve EEA/CH/UK clients with uploaded-video editing needs, or make quality-based purchase decisions before independent benchmarks exist [2][4][5].

The practical takeaway: AI video generation tools 2026 decisions are now cost-per-second decisions — at ~$0.10/sec with no free tier, Gemini Omni 1.1 Flash rewards agencies that build a draft-at-360p workflow [3][5]. Price the service, not the API; the value is the iteration loop around it. For agency pricing and AI adoption context, see how much an AI agency costs in 2026, AI adoption statistics 2026, and our small-business AI tools that actually save money roundup.

9. FAQ: Gemini video API questions for agencies

Is Gemini Omni 1.1 Flash 4K native?

No. 1080p and 4K outputs are upscales of the generated frames, not native renders — Google's docs label them as upscaled [2]. Native generation runs at 720p by default (360p for drafts) [5].

How long can Gemini Omni videos be?

Up to 40 seconds cumulative, built as 10-second extension increments. The model uses up to 10 seconds of prior video as context when extending, so a 30-second ad can be directed as one scene rather than assembled clips [1][2].

What does Gemini Omni cost per second?

Roughly $0.10 per second at 720p through the Gemini API (5,792 tokens per second of 720p video at $17.50 per 1M output tokens) [3]. There is no free tier [3]. 360p drafts cost about one-third as much and render up to 60% faster [1]. Official 1080p/4K rates were not published as of August 2026 [5].

Need the full tool-by-tool breakdown?

See the AI Video Tools Comparison 2026 →

Gemini Omni 1.1 Flash vs Veo 3.1, Sora 2, Kling 3.0, Runway Gen-4.5, and MiniMax H3 — scene length, 4K upscaling, frame control, API access, and use-case fit.

Sources

Notes on accuracy and interpretation