If your H3 Max video keeps zooming when you asked for a moving camera, first decide what would prove the move happened. A larger subject is not enough: a crop can do that. Look for a change in the relationship between the subject and its surroundings.
That gives you something useful to write into the next prompt—and something concrete to check afterward. “More cinematic” gives you neither.
This is a prompting and review guide for AI Live Generator, which runs H3 Max Turbo through its text-to-video and image-to-video endpoints. The prompts below are original starting points, not a set of generated results or a promise of precise camera control. Our current interface offers 5–15 seconds, 480P or 768P, and Balanced or Quality prompt expansion.
Start with the frame that should not move
Consider a red ceramic teapot on a wooden kitchen shelf. A folded linen towel lies nearer to the camera; the tiled wall is farther away. Steam is rising from the spout.
There are three separate kinds of movement here: the steam, the teapot, and the viewpoint. If you want a quiet product shot, only the steam needs to move. You do not have to move the camera to make a still image into a video.
For a text-to-video starting shot, try:
A red ceramic teapot rests on a wooden kitchen shelf. A folded linen towel lies in the foreground, with a pale tiled wall behind the teapot. Medium shot from shelf height, fixed on a tripod. The frame edges and tile lines stay in the same positions throughout. Only a thin thread of steam curls upward from the spout; the teapot stays on the shelf. Soft window light from the left. Audio: quiet kitchen room tone, no speech or music. One uninterrupted shot.The tile lines are deliberate. They give you a visible reference for drift. Watch the edges of the frame as well as the teapot: if the entire scene slowly grows, the camera is not staying locked, even if the subject looks stable in the center.
In image-to-video, replace the scene description with the details that matter in your actual photo. Do not ask for a tiled kitchen if your upload shows a plain studio wall. The image-to-video guide explains how the source frame sets the starting composition.
A push-in needs more than a bigger teapot
In cinematography, a dolly-in changes the camera's position. A lens zoom changes the framing without that position change. In a generated clip, you cannot inspect a physical camera, but you can look for the visual difference: nearer and farther objects should not behave like one flat photograph being enlarged.
Give the scene some depth, then describe the relationship you want to see. Here is an alternative to the locked-shot prompt—not an extra paragraph to append to it:
A red ceramic teapot rests on a wooden kitchen shelf, a folded linen towel in the lower foreground and a tiled wall behind it. The camera glides a short distance forward just above shelf height, ending in a closer view of the teapot. As the viewpoint advances, the nearer towel moves toward the bottom edge faster than the distant tile lines. Keep the teapot on the shelf, with its shape and handle intact. Gentle steam, steady window light. Audio: quiet kitchen room tone, no speech or music. One continuous forward move, no orbit or cut.This asks for depth rather than merely naming a lens. It still does not guarantee correct geometry. If the towel stretches, the handle changes shape, or the wall bends, reject those artifacts even if the overall move looks attractive.
For an uploaded product photo, a modest move is a more conservative request than an extreme close-up. A tightly cropped image may not contain enough visible surroundings to make the requested reveal convincing. Increasing resolution cannot supply a trustworthy view of something the photo never showed.
A sideways move should reveal something
A pan turns the viewpoint from a fixed position. A lateral slide changes its position sideways. To ask for the latter, identify an overlap that should change:
A red ceramic teapot sits on a wooden kitchen shelf. A small stack of recipe cards in the foreground partly covers the lower left side of the teapot; a tiled wall sits behind it. The camera slides gently to the right at shelf height while keeping the teapot in view. During the slide, the nearby cards shift across the frame and uncover the teapot's lower edge. The teapot does not rotate or move along the shelf. Soft stable daylight, a little steam. Audio: quiet kitchen room tone, no music. One short lateral slide with no zoom or cut.The useful question is not “Did it obey the word slide?” It is “Did the overlap change while the teapot kept its shape?” That is a more demanding—and more useful—review.
An orbit is a different request again. It exposes new sides of a subject. If a particular handle, label, or face must remain exact, do not assume a wide orbit will preserve unseen details from one photo. Use a smaller change of viewpoint, or choose a source image that already shows the side you need. See the product video workflow for checking details that must survive the shot.
Pick the next edit from the visible problem
| What you see | What to check before another generation | A focused next change |
|---|---|---|
| The whole picture slowly enlarges | Does the prompt combine “locked” with “push in,” “handheld,” or “reveal”? | Choose one viewpoint behavior and name a background edge that should stay fixed. |
| The subject rotates instead of the camera moving | Does the prompt say “turn,” “rotate,” or “spin” without naming what moves? | Make the camera the subject of the movement sentence; state that the object stays in place. |
| A slide looks like a flat crop | Is there any foreground/background separation to show a changing viewpoint? | Use a frame with visible depth and describe one overlap that should change. |
| The move works but the object deforms | Does the requested viewpoint reveal a large unseen area? | Reduce the move, or replace the input with a better angle. |
| The ending jumps into a new composition | Is a supplied last frame asking for an incompatible viewpoint? | Remove that endpoint for the next diagnostic take, or prepare a compatible one. |
These are possible causes to investigate, not diagnoses you can establish from one failed clip. Keep the input image, duration, resolution, sound instructions, and expansion setting unchanged while you evaluate a wording change. Even unchanged settings do not guarantee identical output, so a better-looking single take is not proof that one edit solved the problem.
The first-and-last-frame guide is useful when the final composition really is a requirement. Two endpoint images do not specify every camera position in between.
Prompt directions are not a camera-path editor
You may have seen other H3 sites with orbit presets, angle values, or draggable camera keyframes. Those controls are not present in AI Live Generator. Here, a camera instruction is part of the prompt; it is not a numerical trajectory.
The Turbo text-to-video API schema documents the prompt and general generation settings. The image-to-video schema also documents first and last image inputs. Neither input list turns a sentence such as “orbit 30 degrees” into an exact camera-path constraint.
Balanced and Quality are prompt-expansion choices, not camera-accuracy levels. Keep one selected while you compare camera wording. Our interface does not offer an expansion-off switch, and a tutorial for another interface should not send you searching for one here.
Set a budget for the question you are answering
A five-second take currently uses 15 credits at 480P or 25 credits at 768P on this site, for either text or image input. Three five-second 480P takes would use 45 credits. Adding one five-second 768P take makes 70 credits in total. These are separate generations, not a free upgrade or an edit to the same file. Check the displayed cost before submitting; package prices are on Pricing.
You do not need to generate all three prompts. Choose the move the shot needs, decide what would count as success, and change the next instruction only when you can say what was wrong with the last result. Stop when another attempt is unlikely to answer a new question.
For more subjects and sound directions, use the H3 Max prompt library. Then bring one camera brief into the generator—with a clear idea of what you will check when the clip comes back.