MiniMax H3 dropped today with open weights, and it’s natively supported in ComfyUI as of this morning. Day zero.
This is a next-generation open-weights video model. Feed it text, images, video, or audio and it generates video with real stereo sound, up to 2K, up to 15 seconds a clip. It is MiniMax’s third-generation video model, following Hailuo 01 and Hailuo 02, and the first the company has released with open weights.
Try on Comfy Cloud
Text-to-video — prompt only.
Image-to-video — bring an image to life.
First-and-last-frame — control the opening frame, the closing frame, or both, and let the model fill in the rest.
Reference-to-video — supply reference images, video, or audio and carry a subject, a motion, or a voice through the clip.
Output runs to 2K and up to 15 seconds. Audio is generated with the video in the same pass, in stereo, not bolted on afterward.
This is the capability MiniMax leads with, and it’s what collapses five separate tasks into one model. Real work rarely draws on one modality. H3 takes images, audio, and video together and resolves them against a prompt that explains how they relate. Describe the relationship between your inputs and the shot you want, and the model handles the cross-modal work itself.
Audio is a property of the model, not a post-process. Every audio output is native stereo.
Motion transfer is the one that matters most for graph work. A reference video can supply movement — a camera move, a performance, a cutting rhythm — while the subject and style come from elsewhere. Combined with in-place editing, that means iterating on a shot.
Getting H3 to run well on consumer hardware took significant machine learning engineering. We found that the model's modulation weights (~40% of the total parameters) could be pruned and replaced with a functionally equivalent lookup table, dramatically shrinking the memory footprint with no loss in output quality.
On top of that, the weights ship with an accurate and efficient int8 convrot quantization, and custom kernels reduce the peak VRAM use during inference.
The result gives a total memory footprint reduced by 66%, from 123.6 GB in full precision to 42.5 GB with the smallest models variants. Combining this with our dynamic VRAM offloading enables a next-generation 2K video model to run locally on a GPU like the RTX 3060.
Update ComfyUI to the latest version 0.30.0 or go to Comfy Cloud
Download the workflows below, or find them in the template library.
Download MiniMax H3 I2V Workflow
Download MiniMax H3 R2V Workflow
Download MiniMax H3 T2V Workflow
Follow the note in the workflow to download the models and save them in the correct model directory.
Write your prompt, connect any frame or reference inputs, and run.
Model weights: 🤗 Comfy-Org/MiniMax-H3
As always, enjoy creating!
No posts




