Gemini Omni 1.1 Flash: What 40-Second 4K Video Means for Agency Content
1. What changed on August 27
On August 27, 2026, Google made Gemini Omni 1.1 Flash generally available through the Gemini video API (model ID gemini-omni-1.1-flash in Google AI Studio), the production successor to the gemini-omni-flash-preview endpoint developers have used since June 30 [1][6]. The headline change for agencies: clips can now be extended in 10-second increments up to a cumulative 40 seconds, with the model using up to 10 seconds of prior video as context — so a 30-second ad or a looping brand asset can be directed as one scene instead of assembled from separate clips [1][2]. The rest covers the three capabilities that matter, Gemini Omni pricing per second, and where it fits in your AI video tools for agencies stack alongside our Gemini model pricing coverage.
2. The three capabilities that matter
1. 40-second scenes with a 10-second lookback. Previous Omni previews referenced only the final second when extending; Omni 1.1 Flash analyzes up to 10 seconds of prior context for visual consistency and narrative flow [2]. A 40 second AI video is now a realistic deliverable — built as 10s increments rather than a single long generation [1][2].
2. First/last frame control. Specify a starting and ending frame, and the model generates continuous video between the keyframes — the primitive behind camera orbits, zoom transitions, and seamless loops [1][2]. Prompt tags <FIRST_FRAME>/<LAST_FRAME> bind images to those roles [2].
3. A 360p draft → 4K upscale ladder. Output resolution runs 360p (draft) → 720p (default) → 1080p (upscaled) → 4K (upscaled) [2][5]. Important accuracy point: 1080p and 4K are upscales of the generated frames, not native renders — Google's docs label them as "upscaled" [2][4]. If you are evaluating a 4K AI video generator, "4K" means a 720p render upscaled [5]. Draft-tier economics (§3) make the workflow clear: iterate at 360p, render the keeper at 720p or upscaled 1080p/4K [1][5].
3. What Gemini Omni 1.1 Flash costs
Gemini Omni pricing is token-based, billed through the Gemini API: input at $1.50 per 1M tokens, text output at $9.00 per 1M, video output at $17.50 per 1M [3]. Video bills at 5,792 tokens per second of 720p — roughly $0.10 per second — so the AI video generation cost per second is the planning unit for client budgets, and a 40-second 720p scene runs about $4.00–$4.06 in output [3]. There is no free tier [3]. A 10-second 360p draft costs ~$0.30–$0.34 versus ~$1.01 at 720p, because 360p renders up to 60% faster at about one-third the cost [1][5].
| Tier | Resolution | Price per 10s clip | Price per 40s scene | Notes | Source |
|---|---|---|---|---|---|
| Draft | 360p | ~$0.30–$0.34 | ~$1.20–$1.36 | up to ~60% faster than 720p, ~1/3 cost | [1][5] |
| Standard | 720p | ~$1.01 | ~$4.00–$4.06 | default; official per-second rate ~$0.10/s | [3] |
| Upscaled | 1080p | ~$1.50 (reseller est.) | ~$6.00 (reseller est.) | official rate not yet published by Google | [5] |
| Upscaled | 4K | ~$3.00 (reseller est.) | ~$12.00 (reseller est.) | official rate not yet published by Google | [5] |
1080p/4K per-second rates are marketplace/reseller estimates (~$0.15/s and ~$0.30/s respectively), not Google's published numbers [5]. Per-clip totals are arithmetic on the cited rates (analysis, not new sourced claims).
The 360p tier is where the economics get interesting: Apidog's worked example puts a 12-draft + 1-final-render workflow at $13.18 if every frame renders at 720p versus $5.07 if drafts run at 360p and only the keeper renders at 720p — about 62% less spend [5]. For context on what agencies charge for AI work overall, see how much an AI agency costs in 2026, and for video-as-a-service retainers, our AI agency pricing negotiation strategies.
4. The Gemini video API for agencies
For agencies building client pipelines, the API story is the practical part. Editing and generation run through the Interactions API — a conversational, multi-turn editing loop where each request references the previous previous_interaction_id, and response_format pins resolution and aspect ratio (9:16 or 16:9) [1][2]. Multimodal input (text, image, audio, video in one call), video references up to 3 clips × 3s each, subject/character reference images, text-in-video rendering, timecode prompting, and native synchronized audio on every clip are all supported [2].
Surfaces: the Gemini API and Gemini Enterprise Agent Platform (Adobe, Figma Weave, GMI Cloud, Runway); consumer-side, Google Flow carries the full feature set for AI Plus/Pro/Ultra subscribers, and scene extension is live in the Gemini app [1][4]. On API quality alone it's a genuine text to video API 2026 contender for short-form work — see our AI workflow automation implementation guide for slotting model APIs into client pipelines.
5. Ads and social video: the workflow change
This is where the update changes a recommendation. A 30-second ad — the most common paid-social creative length — now fits in one directed scene: a beginning frame, an ending frame, and the model fills the middle [1][2]. Combined with the 360p draft tier, the client iteration loop becomes: storyboard and test hooks at 360p (~$0.30–$0.34 per 10s draft), then render the keeper at 720p or upscaled 1080p/4K [1][5]. For ad testing that previously meant five-figure production budgets, a creative variant now costs under a dollar of API spend.
That economics shift is why, for best AI video model for ads decisions, Gemini Omni 1.1 Flash is now the default recommendation for volume and iteration — with the §7 caveat that Veo remains the fidelity pick. For the paid-social angle, our ChatGPT ads for AI agencies piece covers adjacent ad-creative workflows, and AI automation ROI for small business frames justifying the tooling cost to clients.
6. Where it still falls short
Vendor claims are strong; independent proof is not. No third-party benchmarks exist yet for Gemini Omni 1.1 Flash on any public video leaderboard — speed and quality figures are vendor-reported [4].
Operational limits that matter for client work [2]:
- Extension is end-of-clip only — no prepend or mid-clip insertion.
- Uploaded videos for editing/extension must be ≤10 seconds, and EEA, Switzerland, and UK cannot edit or extend uploaded videos at all (model-generated multi-turn editing is supported; uploaded-video workflows are not).
- No adding dialogue to uploaded-video extensions, no voice editing, and audio references unsupported.
- No multi-video referencing, no provisioned throughput, no system instructions/temperature/top_p/negative prompts.
- Every output carries a SynthID watermark; English is fully supported, other languages unevaluated [2].
For agencies selling 4K AI video generator capability to clients, the upscaled-4K framing matters: an upscale, not native-4K rendering [2][5].
7. Veo 3.1 vs Gemini Omni: which to recommend
Google's docs position Gemini Omni Flash as the recommended default for video generation, with Veo as the higher-fidelity specialist — complementary tiers on the same API key, not rivals [2][5]. The practical split: Omni for volume, iteration, and conversational editing; Veo 3.1 for maximum fidelity or longer runtime (up to 148 seconds) [5].
| Capability | Gemini Omni 1.1 Flash | Veo 3.1 |
|---|---|---|
| Price @720p | ~$0.10/s [3] | ~$0.40/s [5] |
| Price @4K | pending (reseller est. ~$0.30/s) [5] | ~$0.60/s [5] |
| Max scene length | 40s cumulative (10s increments) [1][2] | up to 148s [5] |
| Native audio | yes [2] | yes [5] |
| Conversational editing (Interactions API) | yes [1][2] | not stated |
| Free tier | no [3] | not stated |
| Best fit | ads/social short-form, product demos, loops, client iteration | max fidelity + long runtime |
Also on the same key: Veo 3.1 Fast at $0.10/s ($0.30/s @4K) and Veo 3.1 Lite at $0.05/s (no 4K) [5]. If the deliverable is 15–30 second social creatives, Omni's draft-to-final economics win; for long-form brand film, Veo 3.1's 148-second ceiling beats Omni's 40-second cumulative cap. Recommendation update: for ads/social video, Gemini Omni 1.1 Flash becomes the default for volume + iteration; keep Veo 3.1 where max fidelity and >40s runtime are required [2][5]. A full row-by-row breakdown lands on the AI video tools comparison page.
8. Bottom line for agencies
Switch now if you produce short-form social ads, product demos, and looping brand assets — the 40-second ceiling plus 360p drafting changes the cost math for iteration-heavy creative work [1][5]. Wait if you sell long-form brand film (Veo 3.1's 148s ceiling is still the fit), serve EEA/CH/UK clients with uploaded-video editing needs, or make quality-based purchase decisions before independent benchmarks exist [2][4][5].
The practical takeaway: AI video generation tools 2026 decisions are now cost-per-second decisions — at ~$0.10/sec with no free tier, Gemini Omni 1.1 Flash rewards agencies that build a draft-at-360p workflow [3][5]. Price the service, not the API; the value is the iteration loop around it. For agency pricing and AI adoption context, see how much an AI agency costs in 2026, AI adoption statistics 2026, and our small-business AI tools that actually save money roundup.
9. FAQ: Gemini video API questions for agencies
Is Gemini Omni 1.1 Flash 4K native?
No. 1080p and 4K outputs are upscales of the generated frames, not native renders — Google's docs label them as upscaled [2]. Native generation runs at 720p by default (360p for drafts) [5].
How long can Gemini Omni videos be?
Up to 40 seconds cumulative, built as 10-second extension increments. The model uses up to 10 seconds of prior video as context when extending, so a 30-second ad can be directed as one scene rather than assembled clips [1][2].
What does Gemini Omni cost per second?
Roughly $0.10 per second at 720p through the Gemini API (5,792 tokens per second of 720p video at $17.50 per 1M output tokens) [3]. There is no free tier [3]. 360p drafts cost about one-third as much and render up to 60% faster [1]. Official 1080p/4K rates were not published as of August 2026 [5].
Need the full tool-by-tool breakdown?
See the AI Video Tools Comparison 2026 →Gemini Omni 1.1 Flash vs Veo 3.1, Sora 2, Kling 3.0, Runway Gen-4.5, and MiniMax H3 — scene length, 4K upscaling, frame control, API access, and use-case fit.
Sources
- Google blog — "Build with Gemini Omni 1.1 Flash" (Aug 27, 2026): blog.google/…/build-with-gemini-omni-1-1-flash
- Google Gemini API docs — Omni (video generation, extensions, limits): ai.google.dev/gemini-api/docs/omni
- Google Gemini API — pricing (Wayback snapshot 2026-08-27): ai.google.dev/gemini-api/docs/pricing
- OrcaRouter — "Gemini Omni 1.1 Flash launch" (Aug 27, 2026): orcarouter.ai/blog/gemini-omni-1-1-flash-launch
- Apidog — "Gemini Omni 1.1 Flash pricing analysis": apidog.com/blog/gemini-omni-1-1-flash-pricing
- TLDR AI, Aug 28, 2026 issue (GA confirmation): tldr.tech/ai/2026-08-28
Notes on accuracy and interpretation
- 4K/1080p are UPSCALED, not native renders — never read "4K" as native generation; Google's docs label these tiers "upscaled".
- $0.10/s is the 720p standard tier; official 1080p/4K rates are not published — the ~$0.15/s and ~$0.30/s figures are reseller/marketplace estimates.
- "40 seconds" is cumulative via 10s extensions, not a single-shot generation; each extension uses up to 10s of prior video as context.
- No independent third-party benchmarks exist — speed/quality figures are vendor-reported; avoid quality rankings.
- 360p draft "60% faster / ~1/3 cost" figures are Google-reported.
- EEA/CH/UK cannot edit or extend uploaded videos; uploaded-video extensions ≤10s; no voice editing.
- Per-clip cost-math figures are arithmetic on the cited list rates (analysis, not new sourced claims).