H3 Max has a simple strength: it turns a focused shot brief into a short video quickly. That simplicity also gives it clear boundaries. This page separates three things that are often mixed together in search results: the H3 Max model family, the endpoints currently published by fal, and the controls available in this website.
| Question | What this generator supports |
|---|---|
| Native resolution | 480P or 768P |
| Clip duration | 5–15 seconds in whole-second steps |
| Text to video | Yes |
| Image to video | Yes, with a required first frame |
| First and last frame | Yes, with an optional ending image |
| Middle keyframe | No control in this generator |
| Reference video or audio | Not exposed in this generator |
| Native 1080P, 2K, or 4K | Not available in this workflow |
| Downloadable H3 Max weights | Not provided by this website; do not confuse them with the open MiniMax H3 base weights |
| Credits | 1 per second at 480P; 2 per second at 768P |
These are product facts, not plan restrictions. Buying a larger pack gives you more generations; it does not unlock a hidden 2K switch or a longer single clip.
Not in the H3 Max workflow connected to this website. Its native choices are 480P and 768P. The standard MiniMax H3 family has a broader hosted pipeline that includes 2K, which is why 2K claims sometimes appear on pages discussing “H3” without clearly naming the endpoint.
A third-party site may also upscale a 768P result after generation. Upscaling can change output dimensions, but it is not the same as a video generated natively at 2K or 4K. If a page says “H3 Max 4K,” check whether it names the model endpoint and the upscaling step separately.
Use the 480P vs 768P guide to choose the appropriate native setting, or read H3 Max vs MiniMax H3 when 2K is a hard requirement.
This generator accepts a whole-second duration from 5 through 15 seconds. It will reject a 4-second request, a fractional duration, or a single generation longer than 15 seconds before sending it upstream.
For a longer story, treat 15 seconds as a shot boundary rather than asking one generation to do everything. Write separate shots with a stable subject description, camera language, palette, and sound direction, then edit the clips together. This usually gives you more control than an overloaded 30-second prompt would.
In the image mode on this site:
That is not the same as conditioning a generation on a library of character images, video movement, or reference audio. fal now publishes a separate H3 Max Reference-to-Video endpoint for richer multimodal conditioning. This website does not currently expose that endpoint, so you will not see controls here that are disconnected from the active backend.
The distinction matters because early H3 Max articles described reference generation as unavailable, while later official fal documentation added it. Use the current endpoint page—not an undated comparison—as the source of truth.
Sources: fal H3 Max image-to-video, fal H3 Max reference-to-video.
MiniMax H3 is presented by fal as an open-weight base model. H3 Max is fal’s post-trained variant of that base. Those two statements do not make the post-trained H3 Max checkpoint interchangeable with the published MiniMax H3 weights.
This site is a hosted browser interface and does not provide a model download. If a repository claims to offer “H3 Max weights,” verify that it links to an official release from the model owner or post-training team, identifies the exact checkpoint, and states the license. A LoRA or distilled community build based on H3 is not automatically the same model as the hosted H3 Max endpoint.
Source: fal MiniMax H3 overview.
Only two settings change the displayed total in this product:
credit cost = duration × resolution rate
480P = 1 credit per output second
768P = 2 credits per output secondText or image mode, aspect ratio, sound direction, a first frame, and an optional last frame do not add a separate product credit fee. The exact total is printed on the Generate button.
Credits are reserved when a provider job starts. If the provider confirms a terminal failure or policy refusal, this application returns the reserved credits automatically. A successful but creatively disappointing result is still a completed generation, so budget for real iteration rather than assuming every idea will work on the first attempt.
See the H3 Max cost guide for the full duration table and the pricing page for current packs and plans.
An inference timer measures the model-running portion. The time from clicking Generate until a saved MP4 is playable can also include prompt expansion, uploads, provider queueing, webhook or polling delay, durable storage, and network transfer.
We do not publish a site-wide median or P95 until we have a stable, auditable sample. That is more useful than repeating a provider benchmark as if it were a guaranteed delivery time. Read how fast H3 Max is for the two-clock explanation.
The boundaries do not make H3 Max a weak tool. They make it a focused one. It is a practical fit for:
The strongest workflow is usually to solve movement and timing at 480P, then move the selected direction to 768P. If the final deliverable truly needs native 2K, multimodal reference conditioning, existing-video editing, or local model research, choose the broader workflow rather than forcing this one past its design.
No. It is the highest native resolution exposed by this H3 Max generator. A larger plan adds credits, not a higher-resolution mode.
Yes. The first image is required in image mode and the ending image is optional. Prepare compatible compositions and describe the transition between them.
Not currently. fal documents that endpoint, but this product currently connects text-to-video and image-to-video with an optional ending frame.
No. A seed can provide a more stable starting point for comparison, but prompt expansion and generative sampling mean it should not be treated as a frame-for-frame guarantee.
The job initially reserves credits. When the upstream provider confirms the refusal as a terminal failure, the application returns those credits and asks you to revise the input.