Feedback
AI Ad Video Example
Loading...
comfyui minimax h3
Explore comfyui minimax h3: an open-weight ComfyUI setup that renders 2K clips with matching stereo sound from text, stills, or reference media.
All Tools
Discover our comprehensive AI-powered animation toolkit
MiniMax H3
MiniMax H3 AI Video Generator

Seedance 2.0
The Future of AI Video Is Here.

Kling 3.0
Next-Gen AI Video Generator
Grok Video Generator
Create Videos from Text or Images with AI
MiniMax H3 video generator
MiniMax H3 AI Video Generator

Seedance 2.0
The Future of AI Video Is Here.

Kling Motion Control
Turn reference images into amazing motion videos in minutes

Suno AI Music Generator
Create Professional Music with AI
Why Creators Choose the comfyui minimax h3 Workflow
Packed as open weights and wired natively into ComfyUI, the comfyui minimax h3 setup brings MiniMax's omni-modal engine onto your own machine. Text, stills, footage, and audio are read together in one context, so the finished clip arrives with its own stereo soundtrack — dialogue, effects, and score shaped in a single forward pass. Expect roughly 15 seconds of footage at up to 2K and 24fps, with every knob exposed at the node level.
- Audio Generated In-StepDialogue, effects, and music are modeled alongside the picture, so a single pass through the comfyui minimax h3 graph yields one MP4 with everything already in sync.
- Open Weights, Local RunsBecause the comfyui minimax h3 model ships as open weights, you can run it on your own hardware and tune duration, resolution, and diffusion settings without hitting an API ceiling.
- Mix Every Reference TypeFeed text, pictures, footage, and voice samples into the same job — the comfyui minimax h3 nodes hold a character, look, movement, camera angle, or timbre steady across the take.
Three Steps to Your First comfyui minimax h3 Render
Three quick steps take you from a fresh ComfyUI install to open-weight footage with built-in sound via comfyui minimax h3.
Core Capabilities of the comfyui minimax h3 Workflow
Ship-ready templates, omni-modal reasoning, in-pass stereo audio, reference locking, and an optional Sage Attention boost — together the comfyui minimax h3 workflow covers the whole local video pipeline.
Three Ready-Made Templates
Text-to-video, image-to-video, and reference-to-video ship as separate examples in the comfyui minimax h3 library, so each generation mode works the moment you open it.
One Shared Context Window
Rather than bolting modalities together, the comfyui minimax h3 model reads prompts, pictures, footage, and sound in a single context and blends every reference into one render.
Lock Down What Matters
Hold a face, an art style, a motion path, a camera move, or a voice steady using reference media — the comfyui minimax h3 R2V node accepts as many as 9 images, 3 videos, and 3 audio clips.
Clean On-Screen Text and Logos
Lettering and brand marks come out crisp thanks to the comfyui minimax h3 model, whose instruction following lets you spell out how each reference should relate in plain language.
Faster Runs with Sage Attention
Slot the Patch Sage Attention KJ node into your comfyui minimax h3 graph and generation time roughly halves, with barely any visible drop in quality.
Flexible Resolution and Length Grid
The comfyui minimax h3 Resolution Selector derives width and height from your aspect ratio and megapixel target, snapped to a 32-multiple grid and 17-frame blocks at 24fps.
comfyui minimax h3 — Frequently Asked Questions
Answers to the questions people ask most about running MiniMax H3 as a comfyui minimax h3 workflow.
What exactly is the comfyui minimax h3 workflow?
It is the native ComfyUI implementation of MiniMax H3, an omni-modal generation model that MiniMax released with open weights. A single forward pass reads your text, images, video, and audio inputs and returns a clip with its own stereo soundtrack.
How high can the output resolution go?
Expect up to 2K at 24fps across roughly 15 seconds of footage. The native canvas keeps a 768px short edge, tops out at 768x1344, and rounds dimensions to multiples of 32.
Which generation modes ship with it?
Three examples come bundled: text-to-video (T2V), image-to-video (I2V) with optional first and last frame control, and reference-to-video (R2V) for pinning a character, style, motion, camera angle, or voice.
Does the model create audio as well?
It does. Voice, sound effects, and music are all modeled natively in stereo alongside the picture, then delivered together inside one synchronized MP4.
How do I get started?
Move ComfyUI to 0.30.0 or higher, browse Template Library > Video, select a comfyui minimax h3 example, and let the pop-up pull the weights from the Comfy-Org/MiniMax-H3 repository on Hugging Face.
Is there a way to make generation faster?
Install SageAttention along with the KJNodes custom nodes, then insert a Patch Sage Attention KJ node between the UNETLoader and BasicGuider in your comfyui minimax h3 graph — generation time drops by roughly half.
Start Rendering with the comfyui minimax h3 Workflow
Keep everything on your own machine: open weights, stereo sound baked in, and every parameter exposed. Text-to-video, image-to-video, and reference-to-video graphs are ready to load.
