MiniMax H3 is a multimodal AI video model for creators who want more control than a text-only workflow can provide. It can generate video from text, starting images, ending images, reference images, reference clips, and reference audio. In CanvasArc, MiniMax H3 is connected through the official MiniMax API and appears as MiniMax Official in the provider selector.
CanvasArc is an independent creative platform. MiniMax and MiniMax H3 are names and products of their respective owner. CanvasArc is not affiliated with or endorsed by MiniMax.
What is MiniMax H3?
MiniMax H3 is designed for text-to-video, image-to-video, first-and-last-frame generation, and multimodal reference workflows. The official API currently supports:
- 2K video output
- 4–15 second duration
- Native stereo audio
- Prompts up to 7,000 characters
- Common landscape, portrait, square, and adaptive aspect ratios
- Up to 9 reference images
- Up to 3 reference videos and 3 reference audio files
- Up to 12 reference materials in one request
These capabilities make H3 useful when a prompt alone is not enough to preserve a character, reproduce a movement, guide a camera path, or follow the rhythm of an audio reference.
MiniMax H3 generation modes
Text to video
Use text-to-video when the scene can be described without a visual reference. A strong prompt should identify the subject, action, environment, composition, camera movement, lighting, pacing, and desired sound.
Image to video
Use an image as the first frame when you need a specific character, product, composition, or art direction. You can also provide an ending frame to control where the shot finishes.
Multimodal reference video
The H3 reference workflow can combine images, videos, audio, and text. References can guide identity, clothing, visual style, motion, choreography, camera behavior, pacing, or sound. Describe the purpose of every reference instead of uploading files without instructions.
A practical MiniMax H3 prompt formula
Use this structure as a starting point:
[Shot and subject] + [main action] + [environment] +
[camera movement] + [lighting and visual style] +
[timing] + [audio] + [reference instructions]Example:
A cinematic medium shot of a cyclist waiting under a neon-lit awning
during heavy rain. The cyclist looks toward the street as a tram passes.
The camera slowly pushes in, then makes a subtle orbit to the right.
Reflections ripple across the wet pavement, cool blue and magenta light,
realistic skin and fabric detail. Build tension for the first three seconds,
then reveal the tram. Native rain ambience, distant traffic, and a soft
electronic pulse. Use the reference image for the character and clothing.How to generate a MiniMax H3 video in CanvasArc
- Open the CanvasArc AI Video Generator.
- Choose MiniMax Official as the provider.
- Select MiniMax H3 as the model.
- Choose Text to Video, Image to Video, or Video to Video.
- Add the relevant reference media.
- Set a duration between 4 and 15 seconds and select an aspect ratio.
- Write a precise prompt explaining both the desired output and the role of each reference.
- Start the task. You can leave the page and follow its progress from My Creations.
Completed H3 videos are copied to the configured CanvasArc object storage when custom storage is enabled, rather than relying only on a temporary provider URL.
When should you choose MiniMax H3?
MiniMax H3 is a strong choice when you need:
- 2K output with generated stereo audio
- A longer prompt for detailed direction
- First-frame or first-and-last-frame control
- Several visual references in one generation
- Video or audio references for movement and timing
- A task that remains visible in your personal creation history
Other models may be a better fit for a different visual style, provider availability, speed, or price. CanvasArc keeps model names and providers visible so you can choose the actual model used for each task.
MiniMax H3 prompt tips
- Give every uploaded reference a clear job.
- Describe movement and camera behavior with verbs.
- Separate the timeline into beats for complex shots.
- State what must stay consistent, such as face, clothing, product shape, or typography.
- Keep instructions compatible; conflicting references reduce control.
- Use audio direction intentionally: ambience, dialogue, effects, music, rhythm, and timing.
Frequently asked questions
Is MiniMax H3 the same as a Veo model?
No. MiniMax H3 is requested through the official MiniMax API in CanvasArc. Veo models are listed separately with their actual provider and model name.
Does MiniMax H3 support image-to-video?
Yes. It supports a starting frame and can also use a final frame for additional shot control.
Can I close the page while a video is generating?
Yes. CanvasArc stores the generation as a task. Open My Creations to find pending, processing, completed, or failed image and video tasks.
Where can I find the official API specification?
See the official MiniMax video generation documentation for the current API specification and limits.

