Narrative shots
Control how a scene begins, develops, and lands on its final composition.
MiniMax multimodal video
MiniMax's multimodal video model for text-to-video, first- and last-frame animation, and image, video, or audio reference guidance.
From frames to finished motion
Boundary frames establish the composition, the prompt directs the action, and H3 creates the motion between them. This makes the generation path easier to understand before you spend credits.
Try this workflowUpload a first frame, a last frame, or both to define where the shot begins and lands.
Tell H3 what changes between the frames, including subject action, camera movement, timing, and sound.
“The car revs through the desert.”
H3 builds a continuous clip while preserving the intended opening and ending composition.
Generate 4–15 second video at 768P or 2K from text, first and last frames, or image, video, and audio references.
Create with MiniMax H3Control how a scene begins, develops, and lands on its final composition.
Use reference images to guide recurring subjects and visual identity.
Use video and audio references to shape motion, camera language, voice, and pace.
Animate an opening image, land on a final image, or use both to define the arc of a shot while H3 builds the motion between them.
“The car revs through the desert.”
Assign images to character or style, video to motion or camera language, and audio to voice or editing rhythm. H3 accepts up to 9 images, 3 videos, and 3 audio clips.
Explore timing at 768P, then render the selected direction at 2K. Choose any duration from 4 to 15 seconds and see the credit estimate before generation.
Start with the subject, then add lighting, style, and anything that must stay the same. Upload a reference whenever you want a more controlled edit.
Explain whether a reference controls character, motion, camera, style, sound, or timing.
Choose compatible first and last frames, then describe the transition between them.
Validate motion and composition at lower cost before generating the final 2K version.
Choose frontier multimodal control, high-fidelity generation, or a faster production variant.
| Feature | MiniMax H3Current | ||||||
|---|---|---|---|---|---|---|---|
| Best for | Multimodal cinematic direction | Fast, reference-led production | Maximum-speed iteration | MiniMax's high-fidelity video model for realistic human motion, cinematic effects, and expressive character performance. | A faster, lower-cost Hailuo 2.3 variant for image-to-video iteration and production volume. | MiniMax's efficient video model with text-to-video, image animation, and optional first-to-last-frame control. | A low-cost 512p Hailuo 02 variant for rapid concepts with optional first and last frame guidance. |
| Input | Text, image, video, or audio references | Text, frames, image, video, or audio references | Text or first and last frames | Text to video · Image to video | Image to video | Text to video · Image to video | Text to video · Image to video |
| Resolution | 768P or 2K | 480P, 768P, or 1080P | 480P, 768P, or 1080P | 768p or 1080p | 768p or 1080p | 512p, 768p, or 1080p | 512p |
| Frame control | First and last frames | First and last frames | First and last frames | First frame | First frame | First frame · Last frame | First frame · Last frame |
| Speed | Flagship | Faster than real time | Fastest | Hailuo 2.3 | Hailuo 2.3 Fast | Hailuo 02 | Hailuo 02 Fast |
Explore popular image models outside this series, or browse the full catalog.




Start from text, boundary frames, or multimodal references and review the credit estimate before generating.