Foundation video model

Gemini Omni Flash 1.1

Google's stable Gemini Omni Flash 1.1 supports text-to-video, image-to-video, reference-to-video, video editing and extension through the Interactions API, with native audio. Google documents 720p as the default/native render tier and labels 1080p and 4K delivery as upscaled; this page keeps provider aliases and billing units separate from canonical model semantics.

Official site View pricing Updated Sep 22, 2026
Key facts

Verified model snapshot

ReleasedAug 27, 2026Verified Sep 22, 2026
Maximum duration3–10 seconds per generated clip; extension can continue in additional stepsVerified Sep 22, 2026
Maximum resolution4K delivered; 720p native/default; 1080p and 4K upscaledVerified Sep 22, 2026
Native audioYesVerified Sep 22, 2026
API statusavailableVerified Sep 22, 2026
Version delta

What changed

Pricing & providers

Where to access this model

Provider

Google Gemini API

3 observed offers

Last checked Sep 22, 2026

View 3 offer details
Interactions API · all

USD 1.5 / per 1M input tokens

$1.50 per 1M input tokens

Input text, image, video and audio tokens.

Source ↗
Interactions API · video output

USD 17.5 / per 1M video output tokens

$17.50 per 1M video output tokens

Google states 5,792 output tokens per second of 720p video, approximately $0.10/s effective at Standard pricing.

Source ↗
Interactions API · video output · 720p

USD 0.1 / approx effective per second

Approximately $0.10 per 720p output second

Google's stated effective estimate; not a universal resolution tariff.

Source ↗
Provider

fal

16 observed offers · from $0.03/generated second

Last checked Sep 22, 2026

View 16 offer details
google/gemini-omni-flash/v1.1/text-to-video · text to video · 360p

USD 0.03 / generated second

$0.03 per generated second at 360p

fal resolution ladder.

Source ↗
google/gemini-omni-flash/v1.1/text-to-video · text to video · 720p

USD 0.1 / generated second

$0.10 per generated second at 720p

fal resolution ladder.

Source ↗
google/gemini-omni-flash/v1.1/text-to-video · text to video · 1080p

USD 0.15 / generated second

$0.15 per generated second at 1080p

Provider delivery tier; canonical Google docs label 1080p as upscaled.

Source ↗
google/gemini-omni-flash/v1.1/text-to-video · text to video · 4K

USD 0.3 / generated second

$0.30 per generated second at 4K

Provider delivery tier; canonical Google docs label 4K as upscaled.

Source ↗
google/gemini-omni-flash/v1.1/image-to-video · image to video · 360p

USD 0.03 / generated second

$0.03 per generated second at 360p

Source ↗
google/gemini-omni-flash/v1.1/image-to-video · image to video · 720p

USD 0.1 / generated second

$0.10 per generated second at 720p

Source ↗
google/gemini-omni-flash/v1.1/image-to-video · image to video · 1080p

USD 0.15 / generated second

$0.15 per generated second at 1080p

Canonical Google docs label 1080p as upscaled.

Source ↗
google/gemini-omni-flash/v1.1/image-to-video · image to video · 4K

USD 0.3 / generated second

$0.30 per generated second at 4K

Canonical Google docs label 4K as upscaled.

Source ↗
google/gemini-omni-flash/v1.1/reference-to-video · reference to video · 360p

USD 0.03 / generated second

$0.03 per generated second at 360p

Source ↗
google/gemini-omni-flash/v1.1/reference-to-video · reference to video · 720p

USD 0.1 / generated second

$0.10 per generated second at 720p

Source ↗
google/gemini-omni-flash/v1.1/reference-to-video · reference to video · 1080p

USD 0.15 / generated second

$0.15 per generated second at 1080p

Canonical Google docs label 1080p as upscaled.

Source ↗
google/gemini-omni-flash/v1.1/reference-to-video · reference to video · 4K

USD 0.3 / generated second

$0.30 per generated second at 4K

Canonical Google docs label 4K as upscaled.

Source ↗
google/gemini-omni-flash/v1.1/edit-video · video editing · 360p

USD 0.03 / generated second

$0.03 per generated second at 360p

Source ↗
google/gemini-omni-flash/v1.1/edit-video · video editing · 720p

USD 0.1 / generated second

$0.10 per generated second at 720p

Source ↗
google/gemini-omni-flash/v1.1/edit-video · video editing · 1080p

USD 0.15 / generated second

$0.15 per generated second at 1080p

Canonical Google docs label 1080p as upscaled.

Source ↗
google/gemini-omni-flash/v1.1/edit-video · video editing · 4K

USD 0.3 / generated second

$0.30 per generated second at 4K

Canonical Google docs label 4K as upscaled.

Source ↗
Provider

Layer

3 observed offers

Last checked Sep 22, 2026

View 3 offer details
/v2/workspaces/{workspace_id}/base-models/google-gemini-omni-flash-1-1/inferences · video generation · 720p

2 Creative Units per second at 720p

Do not convert Creative Units to USD without plan economics.

Source ↗
/v2/workspaces/{workspace_id}/base-models/google-gemini-omni-flash-1-1/inferences · video generation · 1080p

3 Creative Units per second at 1080p

Do not convert Creative Units to USD without plan economics.

Source ↗
/v2/workspaces/{workspace_id}/base-models/google-gemini-omni-flash-1-1/inferences · video generation · 4K

6 Creative Units per second at 4K

Do not convert Creative Units to USD without plan economics; 4K is provider delivery, not canonical native rendering.

Source ↗
Provider

Cloudflare Workers AI

1 observed offer

Last checked Sep 22, 2026

View 1 offer details
Cloudflare AI model gateway · unified video · 360p / 720p / 1080p / 4K

Availability and schema verified; stable public raw USD price was not captured.

Source ↗
Provider

Runware

1 observed offer

Last checked Sep 22, 2026

View 1 offer details
Runware videoInference · videoInference · 360p / 720p / 1080p / 4K

Provider route and modes verified; raw public price remains null.

Source ↗
Provider

Onysoft

3 observed offers

Last checked Sep 22, 2026

View 3 offer details
Onysoft unified AI API · no video input 10s · 720p

USD 0.94 / per generation

$0.94 per generation

Fresh September 21 route; exact output-duration semantics are not verified, so no per-second normalization.

Source ↗
Onysoft unified AI API · with video input · 360p

USD 1.26 / per generation

$1.26 per generation

Fresh September 21 route; provider alias maps to the canonical Google entity.

Source ↗
Onysoft unified AI API · with video input · 720p

USD 1.26 / per generation

$1.26 per generation

Fresh September 21 route; provider alias maps to the canonical Google entity.

Source ↗
View detailed pricing
Derived signals

What stands out

Evidence-backed signals

  • Native audio is positively verified from an entity-supporting source.
  • Maximum duration is recorded as 3–10 seconds per generated clip; extension can continue in additional steps.
  • Official API access is verified as available.
Public test evidence

Recorded test runs

Hands-on testing pending. No quality score or first-hand verdict is shown until a complete raw test record is attached.
External evidence

Independent reports and examples

These observations belong to the named publishers and providers; they are not FrameSignal hands-on tests.

Web · RunDiffusion / Adam Stewart

video edit object attribute

RunDiffusion shows a before/after edit where two cream armchairs are changed to deep forest-green velvet while the room layout, architecture, daylight and original camera motion are requested to remain fixed.

Reported result: Before/after clips show the requested chair upholstery change with the rest of the scene intended to remain fixed.

Prompt: Change the upholstery of the two cream armchairs to deep forest green velvet. Preserve the chairs shape, room layout, all architecture, daylight and original camera motion. Keep the rest of the video unchanged.

Settings: source video: true · workflow: video_edit

Open source ↗
Web · RunDiffusion / Adam Stewart

reference to video camera pass

RunDiffusion shows a university-courtyard reference workflow and generated camera-pass output for a five-second example, with provider-specific reference-video limits documented on the page.

Reported result: The page shows the supplied courtyard reference and generated camera-pass output.

Prompt: Tracking-shot prompt using a university courtyard reference; full source workflow and media shown on page.

Settings: reference image: true · workflow: reference_to_video · duration example seconds: 5

Open source ↗
Google AI Developers Forum · Zuy-hxf-cffg

same prompt same resolution version comparison

A same-prompt 720p community comparison reports more distorted floral and architectural details in Gemini Omni Flash 1.1 than the previous Omni Flash and links outputs for both versions.

Reported result: The author reports a regression in floral and architectural detail compared with the previous Omni Flash.

Prompt: Long cinematic monsoon-valley prompt with porcelain ruins, ornate helmet, sword, roses and instrumental audio (full prompt in thread).

Settings: resolution: 720p · same prompt: true · comparison: Gemini Omni Flash preview

Open source ↗
Video · Google Flow tutorial creator

multi scene character voice consistency

A Flow tutorial reuses one character across multiple scenes, checks same-face and same-voice behavior, identifies breakpoints, and ends with a same-scene Seedance comparison.

Reported result: The tutorial demonstrates character and voice continuity checks and identifies breakpoints in a multi-scene workflow.

Prompt: Creator publishes the Flow workflow and prompts in the tutorial and reuses one character across multiple scenes.

Settings: surface: Google Flow · same character: true · voice consistency: true · comparison: Seedance 2.0

Open source ↗
Web · Aristotto AI

frame pixel correlation

Aristotto reports greater than 0.998 frame correlation across 720p, 1080p and 4K same-prompt outputs and interprets the higher tiers as upscales rather than independent renders.

Reported metric: >0.998 correlation

Reported result: The source reports >0.998 correlation and concludes higher tiers are upscales, consistent with Google documentation.

Prompt: Same prompt rendered across 720p/1080p/4K for frame-correlation analysis.

Settings: resolutions: ["720p","1080p","4K"] · same prompt: true

Open source ↗
Web · MaxVideoAI

text to video workshop scene

MaxVideoAI publishes a full-prompt 16:9 workshop example with six-second duration, native-audio direction and a raw generated output for inspection.

Reported result: Public example exposes the full prompt, settings and generated output for inspection.

Prompt: Wide 16:9 cinematic workshop scene at night with a bicycle mechanic, smooth waist-high arc, ratchet clicks, tire friction and soft rain (full prompt at source).

Settings: aspect ratio: 16:9 · duration seconds: 6 · audio: true

Open source ↗
External benchmarks

Benchmark snapshots

MaxVideoAI editorial scorecard v1.0.0 · Sep 3, 2026

1.1

editorial scorecardoverall score8.3 score 0 to 10
editorial scorecardprompt8.8 score 0 to 10
editorial scorecardvisual8.5 score 0 to 10
editorial scorecardmotion8.2 score 0 to 10
editorial scorecardconsistency7.9 score 0 to 10
editorial scorecardhuman8.1 score 0 to 10
editorial scorecardtext7.6 score 0 to 10
editorial scorecardaudio8.9 score 0 to 10
editorial scorecardsequencing8 score 0 to 10
editorial scorecardcontrol9.2 score 0 to 10
editorial scorecardstability6.8 score 0 to 10
editorial scorecardvalue8.1 score 0 to 10
Open benchmark source ↗
Capabilities

Capability matrix

Text to videoYes
Image to videoYes
Video to videoYes
Native audioYes
Reference imagesYes
Aspect ratios16:9, 9:16
Compare with

Related public models

Alternatives

Related models to evaluate

Developer access

API

StatusavailableVerified Sep 22, 2026
PricingGoogle bills $1.50 per 1M input tokens and $17.50 per 1M video output tokens; Google states 5,792 output tokens/sec at 720p is approximately $0.10 per second effective. fal separately lists $0.03/$0.10/$0.15/$0.30 per generated second for 360p/720p/1080p/4K.Verified Sep 22, 2026
Official accessDocumentation ↗
Open API record
Latest updates

Gemini Omni Flash 1.1 change feed

View all updates
Evidence

Sources and verification