Post-trained multimodal video model

H3 Max

H3 Max is a fal Research post-trained variant built from MiniMax H3 open weights, with hosted text-to-video, image-to-video, reference-to-video, camera-control, and lip-sync routes.

Official site View pricing Updated Sep 22, 2026
Key facts

Verified model snapshot

ReleasedAug 26, 2026Verified Sep 22, 2026
Maximum duration5–15 secondsVerified Sep 22, 2026
Maximum resolution768P native; 1080P latent refinement on base endpointsVerified Sep 22, 2026
Native audioYesVerified Sep 22, 2026
API statusavailableVerified Sep 22, 2026
Version delta

What changed

No structured version event has been recorded beyond the current public model entry.
Pricing & providers

Where to access this model

Provider

fal

9 observed offers · from $0.025/generated second

Last checked Sep 22, 2026

View 9 offer details
minimax/h3-max/text-to-video · text to video · 480P

USD 0.025 / generated second

$0.025 per generated second; 50% promotional rate

Regular rate is $0.05 per second at 480P.

Source ↗
minimax/h3-max/text-to-video · text to video · 768P

USD 0.04 / generated second

$0.04 per generated second; 50% promotional rate

Regular rate is $0.08 per second at 768P.

Source ↗
minimax/h3-max/text-to-video · text to video · 1080P

USD 0.08 / generated second

$0.08 per generated second; 50% promotional rate

1080P is latent refinement from a native 768P source; regular rate is $0.16 per second.

Source ↗
minimax/h3-max/reference-to-video · reference to video

USD 0.08 / generated second

$0.08 per generated output second

Generated-output component only; reference inputs are billed separately by pooled tokens.

Source ↗
minimax/h3-max/reference-to-video · reference to video

USD 0.02 / per 1000 reference tokens over 4096 free

$0.02 per 1,000 reference tokens over 4,096 free

Do not normalize with output-video per-second pricing.

Source ↗
minimax/h3-max/lip-sync/image-to-video · lip sync · 480P

USD 0.05 / generated second

$0.05 per generated second

Source ↗
minimax/h3-max/lip-sync/image-to-video · lip sync · 768P

USD 0.08 / generated second

$0.08 per generated second

Source ↗
minimax/h3-max/lip-sync/image-to-video · lip sync · 1080P

USD 0.16 / generated second

$0.16 per generated second

Source ↗
minimax/h3-max/lip-sync/image-to-video · lip sync · 2K

USD 0.32 / generated second

$0.32 per generated second

Source ↗
Provider

Pixazo

3 observed offers · from $0.025/generated second

Last checked Sep 22, 2026

View 3 offer details
https://gateway.pixazo.ai/minimax-hailuo-h3-max/v1/text-to-video · text to video · 480P

USD 0.025 / generated second

$0.025 per generated second; current promotional table

Current comparable rate; no unique-cheapest claim is made.

Source ↗
https://gateway.pixazo.ai/minimax-hailuo-h3-max/v1/text-to-video · text to video · 768P

USD 0.04 / generated second

$0.04 per generated second; current promotional table

1080P is described as refinement from native 768P.

Source ↗
https://gateway.pixazo.ai/minimax-hailuo-h3-max/v1/text-to-video · text to video · 1080P

USD 0.08 / generated second

$0.08 per generated second; current promotional table

1080P is described as refinement from native 768P.

Source ↗
Provider

Runware

2 observed offers · from $0.025/generated second

Last checked Sep 22, 2026

View 2 offer details
text to video · 480P

USD 0.025 / generated second

$0.025 per generated second; promotion through September 30

Source ↗
text to video · 768P

USD 0.04 / generated second

$0.04 per generated second; promotion through September 30

Source ↗
Provider

Cloudflare Workers AI

1 observed offer

Last checked Sep 22, 2026

View 1 offer details
Workers AI run endpoint · text or image to video · 480P / 768P

Public page routes pricing to the dashboard; no stable raw USD rate captured.

Source ↗
Provider

Layer

6 observed offers

Last checked Sep 22, 2026

View 6 offer details
video generation

0.42 Creative Units per second

Raw Creative Units retained; no USD conversion.

Source ↗
video generation · 768P

0.672 Creative Units per second

Raw Creative Units retained; no USD conversion.

Source ↗
video generation · 1080P

1.344 Creative Units per second

Raw Creative Units retained; no USD conversion.

Source ↗
reference to video

1.67 Creative Units per second

Raw Creative Units retained; no USD conversion.

Source ↗
reference to video · 768P

2.672 Creative Units per second

Raw Creative Units retained; no USD conversion.

Source ↗
reference to video · 1080P

5.344 Creative Units per second

Raw Creative Units retained; no USD conversion.

Source ↗
View detailed pricing
Derived signals

What stands out

Evidence-backed signals

  • Native audio is positively verified from an entity-supporting source.
  • Maximum duration is recorded as 5–15 seconds.
  • Official API access is verified as available.
Public test evidence

Recorded test runs

Hands-on testing pending. No quality score or first-hand verdict is shown until a complete raw test record is attached.
External evidence

Independent reports and examples

These observations belong to the named publishers and providers; they are not FrameSignal hands-on tests.

Web · Anvisha Pai / Voyager

submit to result wait

A 5-second 768P product turntable completed in one attempt. Voyager reports no reroll or retouch in the common nine-clip fixture; the listed cost is an estimate from the published promotional rate.

Reported metric: 6 seconds

Reported result: 5s clip; one attempt; estimated cost $0.20 at the published promotional rate.

Prompt: A clean product shot that holds its shape through a full slow rotation. (Full prompt preserved at source.)

Settings: endpoint: minimax/h3-max/text-to-video · resolution: 768P · duration seconds: 5 · aspect ratio: 16:9 · prompt expansion mode: disabled

Open source ↗
Web · Anvisha Pai / Voyager

submit to result wait

A busy 768P text-to-video scene retained distinct moving people and objects in Voyager's test, while submit-to-result wait was materially longer than the backend inference examples.

Reported metric: 26 seconds

Reported result: Distinct moving people and objects without merging; estimated cost $0.20.

Prompt: A crowded Saturday farmers market seen from slightly above, shoppers with tote bags, a dog on a leash, a busker with a guitar, bunting swaying, late morning light, gentle camera drift.

Settings: endpoint: minimax/h3-max/text-to-video · resolution: 768P · duration seconds: 5 · aspect ratio: 16:9 · prompt expansion mode: disabled

Open source ↗
Web · Anvisha Pai / Voyager

submit to result wait

An image-to-video test at 768P kept the source silhouette recognizable, but Voyager reports that the dripper filled like a cup rather than following the requested physical action.

Reported metric: 6 seconds

Reported result: Output returned at 768x768; source silhouette remained recognizable but the dripper filled like a cup.

Prompt: The coffee dripper slowly fills as coffee drips through it, steam rising, the camera drifting gently around it, the dripper itself unchanged.

Settings: endpoint: minimax/h3-max/image-to-video · resolution: 768P · duration seconds: 5 · prompt expansion mode: disabled

Open source ↗
Web · Anvisha Pai / Voyager

submit to result wait

The camera-control test kept the product recognizable; Voyager cautions that precise orbit accuracy is difficult to establish from the near-symmetric object used.

Reported metric: 6 seconds

Reported result: Product remained recognizable; precise 3D reconstruction was not established.

Prompt: The coffee dripper is rigid and motionless. Only the camera moves. Preserve its shape and materials and the white background.

Settings: endpoint: minimax/h3-max/camera-controls · resolution: 768P · duration seconds: 5 · camera trajectory: [{"time":0,"azimuth":0,"elevation":0,"distance":1},{"time":1,"azimuth":45,"elevation":10,"distance":1}]

Open source ↗
Web · Anvisha Pai / Voyager

submit to result wait

A lip-sync adapter test retained the supplied audio and changed mouth shape. Voyager explicitly did not score synchronization quality by ear.

Reported metric: 11 seconds

Reported result: Estimated cost $0.6895; supplied audio was retained and mouth shape changed.

Prompt: Animate the generated portrait to the supplied synthetic ad-read audio.

Settings: endpoint: minimax/h3-max/lip-sync/image-to-video · resolution: 480P · duration seconds: 14

Open source ↗
Web · Froging AI Editorial Team

reported result

In a 768P mechanical-transformation test, the main assembly was readable and the design survived into wing beating, but the requested fly-past was missed by the late frame.

Reported result: Main assembly was readable and design survived into wing beating, but the requested fly-past was missed by the late frame.

Prompt: One continuous 5-second macro shot in a sunlit watchmaker's workshop... loose brass gears and cobalt-blue plates rapidly snap together into one mechanical hummingbird... then launches forward before arcing past the camera. (Full prompt preserved at source.)

Settings: resolution: 768P · duration seconds: 5 · aspect ratio: 16:9 · passes: 1 · prompt expansion mode: balanced · safety checker: true

Open source ↗
Web · Froging AI Editorial Team

submit to result wait

A 768P vertical action test measured 6.97 seconds submit-to-result wait while fal inference was reported as 2.79 seconds. Subject and motion remained coherent in sampled frames; fine wheel geometry softened at peak speed.

Reported metric: 6.97 seconds

Reported result: fal inference 2.79s; measured wait 6.97s. Subject and motion remained coherent; fine wheel geometry softened at peak speed.

Prompt: One continuous 5-second vertical documentary shot on a rainy city side street at blue hour... bicycle messenger ... sharp corner ... puddle ... accelerates out of frame. (Full prompt preserved at source.)

Settings: resolution: 768P · duration seconds: 5 · aspect ratio: 9:16 · passes: 1 · prompt expansion mode: balanced · safety checker: true

Open source ↗
Web · Froging AI Editorial Team

submit to result wait

A 768P square stop-motion paper-craft test measured 7.29 seconds submit-to-result wait while fal inference was reported as 1.28 seconds. The broad transformation succeeded and NORTH was legible, but the opening form was not exact.

Reported metric: 7.29 seconds

Reported result: fal inference 1.28s; measured wait 7.29s. Broad transformation succeeded and NORTH was legible; opening form was not exact.

Prompt: One continuous 5-second stop-motion paper-craft shot... folded cobalt-and-coral transit map opens ... six miniature buildings unfold ... sign with the single word NORTH. (Full prompt preserved at source.)

Settings: resolution: 768P · duration seconds: 5 · aspect ratio: 1:1 · passes: 1 · prompt expansion mode: balanced · safety checker: true

Open source ↗
External benchmarks

Benchmark snapshots

Artificial Analysis T2V Leaderboard (With Audio) · Sep 22, 2026

H3 Max

text to videoElo1227 Elo · rank 3 · n=5,689
text to videorank3 rank · rank 3 · n=5,689
text to videosamples5689 comparisons · rank 3 · n=5,689
Open benchmark source ↗
Artificial Analysis I2V Leaderboard (With Audio) · Sep 22, 2026

H3 Max

image to videoElo1195 Elo · rank 1 · n=5,569
image to videorank1 rank · rank 1 · n=5,569
image to videosamples5569 comparisons · rank 1 · n=5,569
Open benchmark source ↗
fal internal human preference evaluation · Sep 22, 2026

H3 Max

mixedoverall preference rank1 rank · rank 1
mixedprompt understanding rank1 rank · rank 1
mixedaesthetics rank1 rank · rank 1
Open benchmark source ↗
Capabilities

Capability matrix

Text to videoYes
Image to videoYes
Native audioYes
Lip syncYes
Reference imagesYes
Aspect ratios21:9, 16:9, 4:3, 1:1, 3:4, 9:16
Compare with

Related public models

Alternatives

Related models to evaluate

Developer access

API

StatusavailableVerified Sep 22, 2026
Official accessDocumentation ↗
Open API record
Latest updates

H3 Max change feed

View all updates
Evidence

Sources and verification