Step 1
Select MiniMax H3
Pick MiniMax H3 in the model selector, then choose a clip length between 5 and 15 seconds and an aspect ratio.
MiniMax H3 is MiniMax's open-weights multimodal video model, out now. Generate native 2K clips at 24 fps, 5 to 15 seconds long, with dialogue, sound effects and ambience produced in the same pass — from text, a still image, or up to nine references.
Prompts run on MiniMax H3 — native 2K with generated audio.
Overview
MiniMax H3 is MiniMax’s next-generation open-weights, general-purpose multimodal video model, released in July 2026 as the successor to Hailuo 2.3 and Hailuo 02. Unlike the Hailuo models before it, H3 takes text, images, video and audio as input and returns native 2K video at 24 fps with a synchronized soundtrack.
A single H3 request generates 5 to 15 seconds of footage in whole-second increments, at 21:9, 16:9, 4:3, 1:1, 3:4 or 9:16 — or in adaptive mode, where the model picks the framing from your reference material. Prompts can run to 7,000 characters, which is enough to direct a multi-shot sequence rather than a single continuous take.
The headline capability is omni-reference. You can attach up to nine reference images, three reference video clips and three reference audio clips to one generation, and H3 will carry character identity, motion and voice across the result. It also accepts instruction-based edits, so an existing clip can be revised in language rather than regenerated from scratch.
MiniMax-H3How It Works
From first prompt to production-ready 2K video in four steps.
Step 1
Pick MiniMax H3 in the model selector, then choose a clip length between 5 and 15 seconds and an aspect ratio.
Step 2
Optional: attach a first or last frame, up to nine reference images, three reference clips, or audio to guide voice and motion.
Step 3
Direct the shot like a filmmaker — subject, camera move, lighting, mood, and the dialogue or sound you want generated.
Step 4
Compare takes, revise with an instruction-based edit instead of restarting, then upscale the one you keep.
Dialogue, sound effects and room tone generated in the same pass as the picture
Native output resolution at 24 fps
Clip length, in whole seconds
Released as an open-weights model by MiniMax
Omni-reference: up to 9 images, 3 videos and 3 audio clips
Animate a still, or lock both ends of the shot
21:9, 16:9, 4:3, 1:1, 3:4, 9:16 — or adaptive
H3 generates dialogue, sound effects and environmental atmosphere alongside the picture, timed to what happens on screen. There is no separate voice or foley step, and no lip-sync pass to line up afterwards — the soundtrack comes out of the same generation as the frames.
Open the generatorAttach up to nine reference images, three reference video clips and three reference audio clips to a single generation. H3 holds character identity, motion style and voice across the shot, so the same person, product or art style survives from one clip to the next.
Open the generatorEvery clip renders at native 2K and 24 fps, up to 15 seconds. Prompts of up to 7,000 characters let you script several shots in one request instead of stitching separate generations, and instruction-based editing revises a finished clip in plain language.
Open the generatorComparison
What changed between MiniMax's previous video model and MiniMax H3.
| Capability | MiniMax H3 | Hailuo 2.3 |
|---|---|---|
| Max resolution | Native 2K | 1080p |
| Frame rate | 24 fps | 24 fps |
| Clip length | 5–15s, whole seconds | 6s at 1080p; 6s or 10s at 768p |
| Native audio | Yes — dialogue, SFX, ambience | No |
| Aspect ratios | 21:9, 16:9, 4:3, 1:1, 3:4, 9:16, adaptive | Not selectable — follows the input |
| Reference images | Up to 9 | Start frame only |
| Reference video | Up to 3 clips (≤15s total) | Not supported |
| Reference audio | Up to 3 clips (≤15s total) | Not supported |
| Instruction-based editing | Yes | No |
| Max prompt length | 7,000 characters | 2,000 characters |
| Weights | Open weights | Closed, API only |
| API model ID | MiniMax-H3 | MiniMax-Hailuo-2.3 |
The hard caps enforced by the MiniMax API on a single generation request.
Prompt Examples
Real prompts behind the example videos on this page — each clip is paired with the exact prompt that generated it. Every prompt opens in the generator running MiniMax H3.
Prompt
GoPro footage of someone airgliding through the mountains, shaky camera footage.
Try this promptGoPro footage of someone airgliding through the mountains, shaky camera footage.
Try this promptA ship sailing through rough waters, grey skies, cinematic ocean spray.
Try this promptPrompt
A silver orb, abstract translucent shape, is floating through the streets of London, shaky camera footage.
Try this promptA silver orb, abstract translucent shape, is floating through the streets of London, shaky camera footage.
Try this promptPrompt
A grassy surface, clear skies, rocky desert, millions of red poppy flowers slowly grow out of the grass.
Try this promptA grassy surface, clear skies, rocky desert, millions of red poppy flowers slowly grow out of the grass.
Try this promptCopy our curated prompts and start creating cinematic AI videos right now.
MiniMax H3 is MiniMax’s next-generation open-weights, general-purpose multimodal video model, released in July 2026. It accepts text, image, video and audio input and generates native 2K video at 24 fps with a synchronized soundtrack. It replaces the Hailuo 2.3 and Hailuo 02 models in MiniMax’s video lineup, and its API model ID is MiniMax-H3.
Yes — they refer to the same model. MiniMax H3 is the official name; because the previous generations were branded Hailuo, the model is also listed as Hailuo 3 on some platforms, and was referred to as Hailuo H3 before launch. MiniMax dropped the Hailuo prefix for this release.
Between 5 and 15 seconds per generation, specified in whole seconds only. That is a 50% increase over Hailuo 2.3, which topped out at 10 seconds and could only reach 1080p at 6 seconds.
Yes. H3 generates dialogue, sound effects and environmental atmosphere in the same pass as the video, timed to on-screen action. Hailuo 2.3 and Hailuo 02 were silent models that needed a separate audio step.
H3 outputs native 2K at 24 fps. You can request 21:9, 16:9, 4:3, 1:1, 3:4 or 9:16, or use adaptive mode and let the model choose the framing based on your reference material.
Omni-reference is H3’s ability to condition one generation on several kinds of reference at once: up to 9 reference images, up to 3 reference video clips, and up to 3 reference audio clips (15 seconds total for each of video and audio). It is how you hold a character’s face, a motion style or a specific voice steady across shots. Reference audio has to accompany an image or video input.
MiniMax released H3 with open weights and describes it as an open general-purpose multimodal video model, which means the weights are published for you to download and run. That is a change from the Hailuo generations, which were API-only.
H3 raises resolution from 1080p to native 2K, extends maximum clip length from 10 to 15 seconds, and adds native synchronized audio, omni-reference conditioning on images, video and audio, multi-shot prompts up to 7,000 characters, and instruction-based editing of existing clips. It is also open-weights, where Hailuo 2.3 was closed.
The API is asynchronous: submit a generation task with model MiniMax-H3 and receive a task_id, poll the task until it succeeds, then fetch the download URL for the finished video. Request bodies are capped at 64 MB, with individual assets limited to 50 MB for video, 30 MB for images and 15 MB for audio.
Type a prompt into the generator at the top of this page and hit Generate — it opens a video tool with MiniMax H3 already selected and your prompt loaded, so you can render a 2K clip with audio without touching the API.
MiniMax's open-weights multimodal video model, out now — native 2K at 24 fps, up to 15 seconds, with dialogue and sound generated alongside the picture.
Try MiniMax H3 Now