Current evidence summary
Google's stable Gemini Omni Flash 1.1 supports text-to-video, image-to-video, reference-to-video, video editing and extension through the Interactions API, with native audio. Google documents 720p as the default/native render tier and labels 1080p and 4K delivery as upscaled; this page keeps provider aliases and billing units separate from canonical model semantics. The source-backed use cases currently point to native-audio text, image and reference-to-video workflows, video editing and extension through the interactions api, teams separating native 720p from delivered/upscaled 1080p and 4k, developers comparing token, per-second, creative unit and per-generation billing.
Public benchmark snapshots
Each result stays attached to its task, date, sample count, and reported uncertainty. Developer backend timings are kept separate from end-to-end wait tests.
1.1
External tests and provider examples
These are reports from the named publishers and providers; FrameSignal did not run these generations.
video edit object attribute
RunDiffusion shows a before/after edit where two cream armchairs are changed to deep forest-green velvet while the room layout, architecture, daylight and original camera motion are requested to remain fixed.
Reported result: Before/after clips show the requested chair upholstery change with the rest of the scene intended to remain fixed.
Prompt: Change the upholstery of the two cream armchairs to deep forest green velvet. Preserve the chairs shape, room layout, all architecture, daylight and original camera motion. Keep the rest of the video unchanged.
Settings: source video: true · workflow: video_edit
Open source ↗reference to video camera pass
RunDiffusion shows a university-courtyard reference workflow and generated camera-pass output for a five-second example, with provider-specific reference-video limits documented on the page.
Reported result: The page shows the supplied courtyard reference and generated camera-pass output.
Prompt: Tracking-shot prompt using a university courtyard reference; full source workflow and media shown on page.
Settings: reference image: true · workflow: reference_to_video · duration example seconds: 5
Open source ↗same prompt same resolution version comparison
A same-prompt 720p community comparison reports more distorted floral and architectural details in Gemini Omni Flash 1.1 than the previous Omni Flash and links outputs for both versions.
Reported result: The author reports a regression in floral and architectural detail compared with the previous Omni Flash.
Prompt: Long cinematic monsoon-valley prompt with porcelain ruins, ornate helmet, sword, roses and instrumental audio (full prompt in thread).
Settings: resolution: 720p · same prompt: true · comparison: Gemini Omni Flash preview
Open source ↗multi scene character voice consistency
A Flow tutorial reuses one character across multiple scenes, checks same-face and same-voice behavior, identifies breakpoints, and ends with a same-scene Seedance comparison.
Reported result: The tutorial demonstrates character and voice continuity checks and identifies breakpoints in a multi-scene workflow.
Prompt: Creator publishes the Flow workflow and prompts in the tutorial and reuses one character across multiple scenes.
Settings: surface: Google Flow · same character: true · voice consistency: true · comparison: Seedance 2.0
Open source ↗resolution upscale frame correlation
Reported metric: frame pixel correlation — >0.998 correlation
Aristotto reports greater than 0.998 frame correlation across 720p, 1080p and 4K same-prompt outputs and interprets the higher tiers as upscales rather than independent renders.
Reported result: The source reports >0.998 correlation and concludes higher tiers are upscales, consistent with Google documentation.
Prompt: Same prompt rendered across 720p/1080p/4K for frame-correlation analysis.
Settings: resolutions: ["720p","1080p","4K"] · same prompt: true
Open source ↗text to video workshop scene
MaxVideoAI publishes a full-prompt 16:9 workshop example with six-second duration, native-audio direction and a raw generated output for inspection.
Reported result: Public example exposes the full prompt, settings and generated output for inspection.
Prompt: Wide 16:9 cinematic workshop scene at night with a bicycle mechanic, smooth waist-high arc, ratchet clicks, tire friction and soft rain (full prompt at source).
Settings: aspect ratio: 16:9 · duration seconds: 6 · audio: true
Open source ↗What official evidence suggests
Promising areas
- Stable GA model with documented T2V, I2V, reference-to-video, edit and extend workflows
- Native audio and 24 FPS with 3–10 second clips
- Google explicitly documents 1080p and 4K as upscaled, preventing a common native-resolution error
- Six current provider surfaces and six accepted external tests are traceable
Verify before paying
- 4K and 1080p are delivered/upscaled outputs rather than independent native renders
- Billing dimensions differ across Google, fal, Layer and Onysoft and should not be collapsed into one cheapest-provider claim
- Reference-video audio is ignored and uploaded audio reference is unsupported
- Third-party evidence includes a dated regression report and is not a FrameSignal hands-on test