If the words have to be correct, do not make H3 Max your typesetter. Generate the shot, leave deliberate space, then add the title, caption, price, disclaimer, or CTA in an editor. Text that appears inside generated video is part of the picture. It can look convincing in one frame and change shape, spelling, or spacing as the clip moves.
This does not mean H3 Max can never produce legible lettering. It means a lucky title card is not a dependable publishing workflow. Exact copy needs a tool that stores characters as characters, not a video model that redraws pixels from frame to frame.
First decide what kind of text you need
“Put text in the video” hides three different jobs:
| Job | Best workflow | Why |
|---|---|---|
| Captions for spoken dialogue | Generate the scene, transcribe the final audio, then add timed captions | The words and timing must match what was actually spoken |
| Headline, price, CTA, or disclaimer | Reserve clean space in the H3 Max shot and typeset the real copy afterward | One wrong character can change the claim or offer |
| A sign, package label, or display that physically belongs in the scene | Start from a prepared image when possible, keep motion conservative, and inspect every frame | The text is part of the object and may deform as the object moves |
There is a fourth case: decorative lettering that does not need to be read. A blurred neon sign or distant menu can be treated as set dressing. Do not confuse that with a readable brand name.
The current fal H3 Max Turbo text-to-video input documents the prompt and generation settings, but no separate subtitle track, caption file, or text-overlay field. AI Live Generator exposes the same practical boundary: you can direct the image and sound, but there is no typography timeline in the generator.
A clean caption area is a visual instruction
“Leave room for captions” is too abstract. Tell the model what should occupy the frame and what should remain quiet.
For a vertical talking shot, that can mean placing the speaker in the upper half, using a simple wall behind the lower third, and keeping hands or props out of the caption zone. For a product clip, it may mean holding the product to one side while the other side stays low contrast.
Here is an original brief for a vertical explainer background:
Vertical 9:16 medium shot of an independent ceramic artist at a workbench.
She lifts one finished cup and turns it once toward the window light. Keep her
face and the cup in the upper two-thirds of frame. The lower third is an
uncluttered warm plaster wall with even light and no objects crossing it.
Locked camera, one continuous take, natural skin and clay texture. No signs,
labels, titles, subtitles, logos, or generated lettering. Audio: one clear line,
“The glaze changes in the kiln,” followed by a soft ceramic tap and quiet room
tone. No music.This prompt is an editorial example, not a claim that a particular generation was tested. Its useful part is the layout: it creates a real surface where accurate captions can be added later. The phrase “no subtitles” reduces the chance that the model invents its own text; it does not create a technical guarantee.
The H3 Max TikTok workflow covers phone-scale framing in more detail. Use it when the final placement is a vertical feed rather than a general video.
Why a clean MiniMax H3 title card is not an H3 Max guarantee
MiniMax H3 and H3 Max are related, but they are not interchangeable output contracts. A guide may show a clean title card from MiniMax H3 at 2K. H3 Max in this generator produces 480P or 768P. Different resolution, model route, prompt, seed, motion, and duration can all change the result.
Even a correctly spelled first frame is not enough. Scrub the full clip and look for:
- a letter changing halfway through a camera move;
- spacing that breathes or crawls;
- a label bending independently of its package;
- a second, invented line appearing nearby;
- text that is technically present but too small for a phone screen.
The H3 Max vs MiniMax H3 comparison explains the larger model distinction. The important editorial rule is simpler: an example proves that one output happened, not that exact typesetting is guaranteed on the next run.
When a prepared first frame helps
If lettering belongs on a physical object, image-to-video gives you a better starting point than asking text-to-video to invent the object and its label together.
Prepare the opening image outside the generator:
- Set the exact label with the real font and approved spelling.
- Make it large enough to survive a 768P frame.
- Keep the object close to its intended hero angle.
- Remove tiny legal copy that viewers could not read anyway.
- Ask for one restrained movement instead of a full spin, splash, orbit, and rack focus.
A suitable motion brief might be:
The bottle remains upright on the stone plinth. Condensation moves slowly down
the glass while a narrow warm reflection travels from left to right. The camera
makes a very slight push in. Preserve the bottle proportions and front label;
do not add, replace, or animate any lettering. Audio: one quiet glass chime and
soft room tone. No music.The prepared frame improves the starting evidence; it does not lock every glyph. Review the label at full size throughout the result. If exact brand fidelity matters more than object motion, generate the environment without the final label and composite the approved packshot afterward.
See the image-to-video guide for source-frame preparation and the product video workflow for conservative motion around packaging.
Captions belong to the final audio, not the prompt
H3 Max can generate dialogue and synchronized sound, but a requested line and the resulting line are not automatically identical. Captioning the prompt before you listen can therefore put the wrong words on screen.
Use this order:
- Generate the video without burned-in captions.
- Listen to the exported audio and decide whether the take is usable.
- Transcribe what is actually spoken.
- Correct names and product terms manually.
- Set caption timing against the final file.
- Export a captioned version and, when the platform supports it, keep a separate caption file for accessibility and later corrections.
If the mouth movement or spoken line is wrong, solve that before captioning. The H3 Max lip-sync troubleshooting guide separates a wording problem from a timing problem.
Use 480P to test layout, 768P for the kept take
On this site, a five-second 480P generation costs 15 credits and a five-second 768P generation costs 25 credits. If you are still deciding whether the subject leaves enough room for a lower third, a 480P layout test can answer that question at lower cost. Once the composition works, use 768P for the take you intend to finish.
Higher resolution gives the editor more pixels for the final overlay. It does not turn model-generated lettering into a guaranteed font engine. The 480P vs 768P guide explains when the extra detail is worth the credit difference.
A final text pass before publishing
Watch the finished export on the smallest screen you expect viewers to use, then check:
- Is every spoken word captioned from the final audio rather than the draft prompt?
- Are product names, prices, dates, units, and disclaimers exact?
- Does the text remain inside the platform's safe area?
- Is there enough contrast behind every line without covering the subject?
- Did any generated sign or label mutate during motion?
- Can you revise the copy without paying to regenerate the video?
That last question is the reason to separate the layers. A price change should require a text edit, not a new AI video. Let H3 Max make the moving picture; keep the words editable until the moment you publish.
Product controls and provider inputs checked October 6, 2026 against this site's current H3 Max implementation and fal's public H3 Max Turbo text-to-video documentation. Model behavior and provider schemas can change.