Under 3 Seconds
Reported by fal for a 5-second, 768p clip. Generating takes less time than video playback, significantly reducing the turnaround between idea and visual draft.
Turn text or an image into a 480p or 768p video with synchronized native audio. H3 Max generates faster than real time, so you can review a shot and explore the next direction sooner.
The H3 Max AI video generator combines fal’s additional training with a jointly optimized inference engine. Its defining advantage is speed, alongside stronger prompt adherence and visual quality than the open-weight base model.
Reported by fal for a 5-second, 768p clip. Generating takes less time than video playback, significantly reducing the turnaround between idea and visual draft.
fal reports about 35× the throughput of the original H3 official endpoint. The comparison is between hosted services; a self-hosted base model’s speed depends on its hardware and inference setup.
Native audio arrives with the video. Include speech, ambience, or action sounds in the prompt, then review the whole scene together as you refine the next take.
From an open-weight base to fal’s post-trained model: the standout change is speed. H3 Max pairs additional training with a purpose-built inference engine to deliver faster results.
Original official endpoint: the baseline in fal’s speed comparison.
5 seconds of 768p video in under 3 seconds; about 35× the baseline throughput, according to fal.
Original H3 model; the base weights provide the starting point for further training.
Additional post-training by fal, co-designed with its own inference engine.
H3 Base weights and inference code are public and support self-hosting. The complete hosted workflow is not included.
Available through hosted APIs. No official public Max weights were found as of September 3, 2026.
Generates video with native audio; provides the foundation for Max.
Retains native audio; fal’s post-training targets stronger prompt adherence and aesthetics alongside speed.
Base: 768p. Official hosted service: 768p or 2K, 4–15 seconds.
480p or 768p, 5–15 seconds. The under-3-second figure applies to a 5-second clip at 768p.
Browse creative directions for your next scene, from product motion to stylized characters. Use these references to shape your own shot brief.
Follow the tutorial to prepare an input, describe one shot, select the available settings, and review your H3 Max video.
For a quick first pass, use 768p, a 5–10 second clip, and Balanced prompt expansion. Describe one scene and its sound, then change one detail per iteration. H3 Max supports 480p or 768p and 5–15 second clips.
Name the person, product, object, or place, plus any clothing, material, color, or scale that must remain recognizable.
Choose one main action. Add smaller movements only when they help the scene, and give each beat enough time to read.
Define the place, time, lighting, weather, and useful background activity to keep the visual context stable.
Choose the framing and one camera move, such as a slow push-in, tracking shot, locked wide shot, or controlled orbit.
Write the opening, change, and ending in order. For two frames, describe the movement that connects them.
Name the dialogue, ambience, music, or action sounds you want and when they should occur.
H3 Max is most useful when one idea needs several attempts. Faster generation leaves more time to compare shots and refine the strongest direction.
Try a product reveal, an action, and a close-up for the same campaign. Compare the first few seconds before spending time on a full edit.
Turn a fresh hook into a short draft, then adjust its framing or timing. A quick feedback loop helps you explore the idea while its direction is still clear.
Try one short line and one visible action. Review whether they explain the idea together, then refine the next version. Short clips can become building blocks for a longer lesson.
Start with a product image and compare a gentle orbit with a detail reveal. Keep checking shape and label accuracy as you narrow down the strongest shot.
Turn a proposed camera move into a short visual draft. Use it to discuss framing and pacing, then try the next direction while the team is still reviewing the scene.
Try the same simple action in clay, paper, or pixel art. Keep the brief comparable so each short generation helps you choose a visual direction.
Choose text-to-video, image-to-video, or first-and-last-frame input based on the material you already have.
Start from language
Start with text for an original scene. Describe the subject, action, camera, sound, and output ratio.
Use for visual concepts, atmosphere studies, or trying different camera directions.
Start from a visual reference
Use an image to define the person, product, or opening composition, then describe what moves next. The video ratio follows the image.
Use for product photos, portraits, illustrations, and a recognizable opening frame.
Guide both ends
Add an ending frame to guide the final pose or composition. Describe plausible motion while keeping the subject and style compatible.
Use for a planned reveal, a pose change, or a shot with a specific closing image.
Match the starting point to the task
Text leaves room to explore, an image anchors the opening, and two frames guide both ends. Always describe the motion between them.
Learn how H3 Max speed, native audio, output settings, and model access differ from the open-weight H3 base.
Your next video
Start with text or an image, generate a short video, and refine the strongest direction in the video workspace.