Feedback
AI Ad Video Example
Loading...
comfyui minimax h3
Run the comfyui minimax h3 node pack inside ComfyUI to turn text, images, or clips into open-weight video with native stereo audio, up to 2K at 24fps.
All Tools
Discover our comprehensive AI-powered animation toolkit

Nano Banana Pro
Advanced AI Image Generator
Sora2 AI
Advanced AI Video Generator for High-Quality Videos

Moive Maker
Turn Ideas into Stunning Movies with AI

Veo3.1 Video Generator
Create Stunning Videos with Veo3.1

3D Science Video
Create 3D science videos easily

AI Ads Video Generator
Generater Ads Video With AI

AI Educational Video Generator
Generate Educational videos easily

Photo to dance Video
AI Dance Video Generator
Inside the comfyui minimax h3 Workflow for Open-Weight Video
The comfyui minimax h3 workflow brings MiniMax's omni-modal generation model into ComfyUI as downloadable open weights. It reads text, images, video, and audio inside one shared context, then renders footage together with stereo sound — dialogue, effects, and music synthesized in the same forward pass. Clips reach 2K at 24fps for roughly 15 seconds, and every diffusion parameter stays adjustable at the node level.
- Built-In Stereo SoundDialogue, effects, and music are synthesized alongside the picture and muxed into one MP4, kept in sync throughout the comfyui minimax h3 run.
- Local Open-Weight RunsBecause the comfyui minimax h3 weights execute on your own machine, resolution, duration, and each diffusion setting stay editable with no API caps.
- Mix Any Reference TypeFeed text, stills, clips, and audio into a single pass to pin down a character, look, motion, camera path, or voice through the comfyui minimax h3 nodes.
Running the comfyui minimax h3 Workflow in Three Steps
Go from a fresh install to open-weight video with native audio in three short steps using the comfyui minimax h3 workflow.
Feature Highlights of the comfyui minimax h3 Workflow
From three ready-made ComfyUI templates to omni-modal understanding, native stereo sound, reference locking, and optional Sage Attention acceleration, the comfyui minimax h3 workflow covers the whole local production chain.
Three Ready-Made Templates
The comfyui minimax h3 template library includes text-to-video, image-to-video, and reference-to-video examples, so every generation mode works straight after install.
One Shared Multimodal Context
Text, stills, footage, and audio are all read inside the same context by the comfyui minimax h3 model, letting any mix of references drive a single render.
Lock Looks From References
Pin a face, art style, motion, camera move, or voice straight from your source files — up to 9 images, 3 videos, and 3 audio clips through the comfyui minimax h3 R2V node.
Clean Text and Logo Rendering
Lettering and brand marks come out legible with the comfyui minimax h3 model, and natural-language instructions let you spell out how each reference relates.
Sage Attention Acceleration
Drop a Patch Sage Attention KJ node into the comfyui minimax h3 workflow to roughly double render speed while barely touching visual quality.
Resolution and Duration Grid
The comfyui minimax h3 Resolution Selector derives width and height from aspect ratio and megapixels, snapped to the model's 32-multiple grid and 17-frame blocks at 24fps.
comfyui minimax h3 — Common Questions
Answers to the questions people ask most about running the MiniMax H3 model locally inside ComfyUI.
What exactly is the comfyui minimax h3 workflow?
It is the native ComfyUI integration of MiniMax H3, an omni-modal generation model MiniMax released with open weights. A single forward pass turns text, images, video, and audio references into video with its own stereo soundtrack.
How good is the output quality?
Renders from the comfyui minimax h3 workflow go up to 2K at 24fps and last about 15 seconds. The native canvas keeps a 768px short edge, tops out at 768x1344, and rounds dimensions to multiples of 32.
Which generation modes ship with it?
Three examples come bundled: text-to-video (T2V), image-to-video (I2V) with optional first and last frame control, and reference-to-video (R2V) for locking a character, style, motion, camera path, or voice.
Does the workflow produce audio as well?
It does. Voice, sound effects, and music are modeled natively in stereo by the comfyui minimax h3 model in the same pass as the picture, then delivered synced inside one MP4.
What do I need to get started?
Update ComfyUI to 0.30.0 or newer, open Template Library > Video, pick a comfyui minimax h3 template, and let the pop-up pull the weights from the Hugging Face Comfy-Org/MiniMax-H3 repository.
Is there a way to speed things up?
Install SageAttention plus the KJNodes pack, then insert a Patch Sage Attention KJ node between the UNETLoader and BasicGuider in the comfyui minimax h3 workflow for roughly double the speed.
Put the comfyui minimax h3 Workflow to Work
Run MiniMax H3 on your own machine with open weights, stereo audio, and full parameter control — text-to-video, image-to-video, and reference-to-video templates are ready when you are.
