Verified model snapshot
What changed
Artificial Analysis' Sep-22 image-to-video-with-audio snapshot records Elo 1099, row rank 8, statistical range 7–9 and 4,861 comparisons.
Current xAI documentation confirms T2V, I2V, R2V, first/last-frame controls, video edit and extension, native audio, up to seven reference images and three preset voices.
Where to access this model
xAI
4 observed offers · from $0.08/generated second
Last checked Sep 22, 2026
View 4 offer details
USD 0.08 / generated second
$0.08 per generated second at 480p
Canonical xAI direct output rate.
Source ↗USD 0.14 / generated second
$0.14 per generated second at 720p
Canonical xAI direct output rate; R2V remains capped at 720p.
Source ↗USD 0.25 / generated second
$0.25 per generated second at 1080p
Canonical xAI direct output rate; native 1080p applies to T2V/I2V.
Source ↗USD 0.01 / input image
$0.01 per input image
Separate image-input charge; do not merge with output-video rate.
Source ↗fal
3 observed offers · from $0.08/generated second
Last checked Sep 22, 2026
View 3 offer details
USD 0.08 / generated second
$0.08 per generated second at 480p
fal route; provider surface may expose a narrower subset than xAI direct.
Source ↗USD 0.14 / generated second
$0.14 per generated second at 720p
fal route; direct xAI and fal nominal rates are kept as separate offers.
Source ↗USD 0.25 / generated second
$0.25 per generated second at 1080p
fal route; provider capability subset remains separate from canonical xAI limits.
Source ↗Roko API
3 observed offers
Last checked Sep 22, 2026
View 3 offer details
USD 0.099 / per generation
$0.099 per generation for 6–10s
Non-official TTAPI route; do not convert to a canonical per-second rate.
Source ↗USD 0.22 / per generation
$0.22 per generation for 12–20s
Non-official TTAPI route; the 20s bucket exceeds canonical xAI max and stays provider-specific.
Source ↗USD 0.33 / per generation
$0.33 per generation for 21–30s
Non-official TTAPI route; do not publish 30s as xAI's canonical limit.
Source ↗reAPI
4 observed offers · from $0.04/generated second
Last checked Sep 22, 2026
View 4 offer details
USD 0.04 / generated second
$0.04 per generated second at 480p
Provider route label preserved; route identity with xAI direct is not assumed.
Source ↗USD 0.08 / generated second
$0.08 per generated second at 720p
Provider route label preserved; provider-specific pricing is not canonical xAI pricing.
Source ↗USD 0.084 / generated second
$0.084 per generated second at 480p
Separate reAPI route label; do not merge with official-route row.
Source ↗USD 0.144 / generated second
$0.144 per generated second at 720p
Separate reAPI route label; do not merge with official-route row.
Source ↗EvoLink
2 observed offers · from $0.056/generated second
Last checked Sep 22, 2026
View 2 offer details
USD 0.056 / generated second
$0.056 per generated second at 480p
Gateway uses preview alias; provider subset remains separate from current xAI canonical surface.
Source ↗USD 0.0992 / generated second
$0.0992 per generated second at 720p
Gateway uses preview alias; provider subset remains separate from current xAI canonical surface.
Source ↗Runware
4 observed offers
Last checked Sep 22, 2026
View 4 offer details
USD 1.5 / per example generation
$1.50 for one 6-second example
Provider example; approximately 3m42s end-to-end, not canonical inference speed.
Source ↗USD 0.99 / per example generation
$0.99 for one 7-second example
Provider example with preset voice; approximately 42s end-to-end.
Source ↗USD 0.73 / per example generation
$0.73 for one 5-second example
Provider example with 3 reference images; approximately 43s end-to-end.
Source ↗USD 1.4 / per example generation
$1.40 for one 10-second example
Provider example with two preset voices; approximately 55s end-to-end.
Source ↗What stands out
Evidence-backed signals
- Native audio is positively verified from an entity-supporting source.
- Maximum duration is recorded as 1–15 seconds canonical xAI generation; longer Roko buckets are provider-specific and non-official.
- Official API access is verified as available.
Objective collection placement
Recorded test runs
Independent reports and examples
These observations belong to the named publishers and providers; they are not FrameSignal hands-on tests.
provider hands on i2v
EvoLink's product-ad I2V test kept the perfume-bottle composition convincing through a restrained orbit and light sweep, while delicate requested audio was limited.
Reported result: Bottle/composition stayed convincing; restrained orbit/light sweep worked, but requested delicate audio was limited.
Prompt: Premium perfume bottle single-shot test; exact prompt preserved at source.
Settings: source image: true · duration seconds: 6 · resolution: 720p · single continuous shot: true
Open source ↗provider hands on talking head i2v
EvoLink's talking-head I2V test produced usable short American-English speech, small head motion and room tone; lip-sync and pronunciation still needed review.
Reported result: Short American-English speech, small head motion and room tone were usable; author says lip-sync/pronunciation still need review.
Prompt: Direct-to-camera English talking-head test; exact prompt preserved at source.
Settings: source image: true · duration seconds: 6 · resolution: 720p · single continuous shot: true
Open source ↗provider hands on cinematic i2v
EvoLink's desert-courier cinematic I2V test delivered strong mood, with local deformation risk when wind affected scarf, coat and hair.
Reported result: Strong cinematic mood; local deformation risk increased when wind affected scarf, coat and hair.
Prompt: Desert courier cinematic test with wind, sand, lightning and thunder; exact prompt preserved at source.
Settings: source image: true · duration seconds: 6 · resolution: 720p · single continuous shot: true
Open source ↗raw generation charge
Runware's Raku pottery workshop text-to-video example shows a $1.50 raw generation charge and approximately 3m42s end-to-end wait.
Reported metric: 1.5 USD
Reported result: Provider displays approximately 3m42s end-to-end and raw output.
Prompt: Raku pottery workshop social ad; full prompt preserved at source.
Settings: duration seconds: 6 · dimensions: 1904x1072
Open source ↗raw generation charge
Runware's first-frame-with-voice example shows a $0.99 raw generation charge and approximately 42 seconds end-to-end.
Reported metric: 0.99 USD
Reported result: Provider displays approximately 42 seconds end-to-end and raw output.
Prompt: Mobile knife sharpening reel; full prompt preserved at source.
Settings: duration seconds: 7 · resolution: 720p · preset voice: rex · first frame: true
Open source ↗raw generation charge
Runware's three-reference logo-loop example shows a $0.73 raw generation charge and approximately 43 seconds end-to-end.
Reported metric: 0.73 USD
Reported result: Provider displays approximately 43 seconds end-to-end and raw output.
Prompt: Avalanche forecast center logo loop; full prompt preserved at source.
Settings: duration seconds: 5 · resolution: 1280x720 · reference images: 3
Open source ↗raw generation charge
Runware's multi-speaker dialogue example shows a $1.40 raw generation charge and approximately 55 seconds end-to-end.
Reported metric: 1.4 USD
Reported result: Provider displays approximately 55 seconds end-to-end and raw output.
Prompt: Oyster hatchery documentary B-roll with two preset voices; full prompt preserved at source.
Settings: duration seconds: 10 · dimensions: 720x1280 · preset voices: ["iris","atlas"]
Open source ↗Benchmark snapshots
1.5
Capability matrix
Related public models
Related models to evaluate
Gemini Omni Flash 1.1
Direct alternative: both are foundation model products with shared text to video, image to video, video to video, cinematic, native audio, character consistency workflows.
Q3 Turbo
Direct alternative: both are foundation model products with shared text to video, image to video, cinematic, native audio, character consistency, social content workflows.
H3 Max
Direct alternative: both are foundation model products with shared text to video, image to video, cinematic, native audio, character consistency workflows.
API
Grok Imagine Video 1.5 change feed
Artificial Analysis' Sep-22 image-to-video-with-audio snapshot records Elo 1099, row rank 8, statistical range 7–9 and 4,861 comparisons.
Current xAI documentation confirms T2V, I2V, R2V, first/last-frame controls, video edit and extension, native audio, up to seven reference images and three preset voices.
Sources and verification
official blog · Official model source · Accessed Sep 22, 2026Grok Imagine 1.5 Preview
official blog · Official model source · Accessed Sep 22, 2026grok-imagine-video-1.5 model card
official docs · Official model source · Accessed Sep 22, 2026xAI Video Generation
official docs · Official model source · Accessed Sep 22, 2026xAI Video gRPC API Reference
official api · Official model source · Accessed Sep 22, 2026xAI API Pricing
official docs · Official model source · Accessed Sep 22, 2026xAI Release Notes
official docs · Official model source · Accessed Sep 22, 2026Roko Grok Imagine Video 1.5 Guide
third party · Background only · Accessed Sep 22, 2026fal Grok Imagine Video 1.5
third party · Background only · Accessed Sep 22, 2026fal Text-to-Video 1.5
third party · Background only · Accessed Sep 22, 2026reAPI Grok Imagine Video 1.5
third party · Background only · Accessed Sep 22, 2026EvoLink Grok Imagine Video 1.5
third party · Background only · Accessed Sep 22, 2026Runware Grok Imagine Video 1.5 Examples
third party · Background only · Accessed Sep 22, 2026Artificial Analysis Image to Video Leaderboard
third party · Background only · Accessed Sep 22, 2026EvoLink Real Samples Review
third party · Background only · Accessed Sep 22, 2026xAI Reference to Video
official docs · Official model source · Accessed Sep 22, 2026