Saturday, September 12, 2026

fal’s H3 Max Generates a 5-Second AI Video With Audio in 1.54 Seconds of Inference

By

Published

3 min read

F

Updated benchmarks confirm that fal’s H3 Max and H3 Max Turbo deliver finished video with synchronized audio in less time than the clips themselves last.

SAN FRANCISCO, CA, UNITED STATES, September 8, 2026 /EINPresswire.com/ — fal today released measured generation times for H3 Max and H3 Max Turbo, its post-trained versions of the open-weights MiniMax H3 video model. Across nine different setups covering all supported resolutions and clip lengths, both endpoints output video faster than the clip plays. H3 Max by fal is the fastest AI video generator of 2026!

The numbers represent generation time—the inference value each endpoint reports on its own response, meaning the model’s work on the GPU. A caller’s total wait also includes queue time, prompt expansion, and encoding on top of that. Measurements were taken on September 8, 2026, using text-to-video with prompt expansion left at its default setting.

At 768p, the resolution for which the models are optimized:

– 5 seconds of video takes 1.54 seconds on H3 Max Turbo and 2.46 seconds on H3 Max

– 10 seconds takes 4.29 seconds on Turbo and 7.55 seconds on H3 Max

– 15 seconds takes 8.44 seconds on Turbo and 15.17 seconds on H3 Max

At 480p, a 5-second clip finishes in 0.44 seconds on Turbo and 0.75 seconds on H3 Max—roughly 11 times and 7 times faster than the clip plays. At 1080p, a 5-second clip requires 2.33 seconds on Turbo and 3.12 seconds on H3 Max. Turbo delivered every one of the nine configurations faster than real time, while H3 Max did so in seven of the nine.

Video Length | Quality | H3 Max Turbo by fal | H3 Max by fal

5 second clip | 480p | .44 seconds | .75 seconds

10 second clip | 480p | 1.00 seconds | 1.72 seconds

15 second clip | 480p | 1.89 seconds | 3.14 seconds

5 second clip | 768p | 1.54 seconds | 2.46 seconds

10 second clip | 768p | 4.29 seconds | 7.55 seconds

15 second clip | 768p | 8.44 seconds | 15.17 seconds

5 second clip | 1080p | 2.33 seconds | 3.12 seconds

10 second clip | 1080p | 6.81 seconds | 8.89 seconds

15 second clip | 1080p | 13.56 seconds | 17.63 seconds

Both picture and audio are produced in a single pass. A default request returns a 5-second 768p clip at 1344 by 768 and 24 frames per second, with stereo audio in the same file. This means dialogue, room tone, effects, and music are described in the prompt rather than added in a second step.

fal credits the speed to designing the model and the inference engine as one unified system, rather than training first and serving later. On the model side, fal Research added new training data and used an in-house reinforcement learning framework to improve prompt adherence and visual quality, keeping optimizations such as lower precision and reduced sampling steps only where human preference scores held.

On quality, Artificial Analysis ranks H3 Max first among image-to-video models with audio output in its Image to Video Arena, at an Elo of 1200—ahead of the next model at 1192 and the base MiniMax H3 model at 1187. Those standings were recorded on September 8, 2026.

H3 Max is available now through fal’s serverless API across six endpoints: text-to-video, image-to-video, and reference-to-video on H3 Max; text-to-video and image-to-video on H3 Max Turbo; and H3 Max Director, a realtime model that maintains a streaming session instead of answering a single request. Clips range from 5 to 15 seconds at 480p, 768p, or 1080p, in aspect ratios from 21:9 to 9:16, with prompts up to 50,000 characters. Python and JavaScript SDKs are available, and the API can be called directly over HTTP.

Pricing is per second of output video with no minimums and no subscription. H3 Max is listed at 0.05 US dollars per second at 480p, 0.08 at 768p, and 0.16 at 1080p. H3 Max Turbo is listed at 0.025, 0.04, and 0.08. Both are running at 75 percent off through September 14, 2026. Signed-in users receive five free generations per day of up to 15 seconds each, on a rolling 24-hour reset.

H3 Max and H3 Max Turbo are available at fal.ai.

About fal

fal is a generative media platform that serves image, video, and audio models through a serverless API, running them on infrastructure it builds and optimizes itself. Learn more at fal.ai.

Bennett Heyn

fal

email us here


David Hall

David Hall

David is the senior editor at FintechNewsWatch. He has a background in journalism and has worked with various media outlets, covering topics ranging from digital banking and blockchain technology to startup funding and regulatory developments. When he is not writing, David enjoys reading, hiking, photography, and exploring new coffee shops.