Current evidence summary
xAI's generally available Grok Imagine Video 1.5 now covers text-to-video, image-to-video, reference-to-video, first/last-frame control, video editing and extension with native generated audio. This page separates the current canonical 15-second xAI limit and native 1080p T2V/I2V surface from provider-specific subsets and Roko's non-official 30-second route. The source-backed use cases currently point to native-audio text, image and reference-to-video workflows, teams needing first/last-frame, edit and extension controls, developers comparing xai direct pricing with provider route billing, current benchmark and preview-to-ga capability tracking.
Public benchmark snapshots
Each result stays attached to its task, date, sample count, and reported uncertainty. Developer backend timings are kept separate from end-to-end wait tests.
1.5
External tests and provider examples
These are reports from the named publishers and providers; FrameSignal did not run these generations.
provider hands on i2v
EvoLink's product-ad I2V test kept the perfume-bottle composition convincing through a restrained orbit and light sweep, while delicate requested audio was limited.
Reported result: Bottle/composition stayed convincing; restrained orbit/light sweep worked, but requested delicate audio was limited.
Prompt: Premium perfume bottle single-shot test; exact prompt preserved at source.
Settings: source image: true · duration seconds: 6 · resolution: 720p · single continuous shot: true
Open source ↗provider hands on talking head i2v
EvoLink's talking-head I2V test produced usable short American-English speech, small head motion and room tone; lip-sync and pronunciation still needed review.
Reported result: Short American-English speech, small head motion and room tone were usable; author says lip-sync/pronunciation still need review.
Prompt: Direct-to-camera English talking-head test; exact prompt preserved at source.
Settings: source image: true · duration seconds: 6 · resolution: 720p · single continuous shot: true
Open source ↗provider hands on cinematic i2v
EvoLink's desert-courier cinematic I2V test delivered strong mood, with local deformation risk when wind affected scarf, coat and hair.
Reported result: Strong cinematic mood; local deformation risk increased when wind affected scarf, coat and hair.
Prompt: Desert courier cinematic test with wind, sand, lightning and thunder; exact prompt preserved at source.
Settings: source image: true · duration seconds: 6 · resolution: 720p · single continuous shot: true
Open source ↗provider documented t2v
Reported metric: raw generation charge — 1.5 USD
Runware's Raku pottery workshop text-to-video example shows a $1.50 raw generation charge and approximately 3m42s end-to-end wait.
Reported result: Provider displays approximately 3m42s end-to-end and raw output.
Prompt: Raku pottery workshop social ad; full prompt preserved at source.
Settings: duration seconds: 6 · dimensions: 1904x1072
Open source ↗provider documented first frame
Reported metric: raw generation charge — 0.99 USD
Runware's first-frame-with-voice example shows a $0.99 raw generation charge and approximately 42 seconds end-to-end.
Reported result: Provider displays approximately 42 seconds end-to-end and raw output.
Prompt: Mobile knife sharpening reel; full prompt preserved at source.
Settings: duration seconds: 7 · resolution: 720p · preset voice: rex · first frame: true
Open source ↗provider documented r2v
Reported metric: raw generation charge — 0.73 USD
Runware's three-reference logo-loop example shows a $0.73 raw generation charge and approximately 43 seconds end-to-end.
Reported result: Provider displays approximately 43 seconds end-to-end and raw output.
Prompt: Avalanche forecast center logo loop; full prompt preserved at source.
Settings: duration seconds: 5 · resolution: 1280x720 · reference images: 3
Open source ↗provider documented dialogue
Reported metric: raw generation charge — 1.4 USD
Runware's multi-speaker dialogue example shows a $1.40 raw generation charge and approximately 55 seconds end-to-end.
Reported result: Provider displays approximately 55 seconds end-to-end and raw output.
Prompt: Oyster hatchery documentary B-roll with two preset voices; full prompt preserved at source.
Settings: duration seconds: 10 · dimensions: 720x1280 · preset voices: ["iris","atlas"]
Open source ↗What official evidence suggests
Promising areas
- Current xAI docs cover T2V, I2V, R2V, first/last-frame, editing and extension
- Native generated audio with up to 7 reference images and 3 preset voices
- Native 1080p T2V/I2V is distinguished from the 720p R2V and edit ceilings
- Seven accepted provider-authored examples preserve prompts, settings, costs or timing
Verify before paying
- Canonical xAI generation remains capped at 15 seconds; Roko's longer route is non-official
- Provider subsets do not expose the full xAI capability surface
- Per-second direct rates and fixed per-generation gateway charges are not directly equivalent
- Current Artificial Analysis snapshot is rank 8 rather than the June launch-era #1 claim