ComfyUI MiniMax H3 Video Generator
Type a scene, drop in a reference, and watch the comfyui minimax h3 graph produce video with matching stereo sound.
AI Video Prompt Generator
10s

Feedback

AI Ad Video Example

Loading...

comfyui minimax h3

Explore comfyui minimax h3: an open-weight ComfyUI setup that renders 2K clips with matching stereo sound from text, stills, or reference media.

All Tools

Discover our comprehensive AI-powered animation toolkit

Why Creators Choose the comfyui minimax h3 Workflow

Packed as open weights and wired natively into ComfyUI, the comfyui minimax h3 setup brings MiniMax's omni-modal engine onto your own machine. Text, stills, footage, and audio are read together in one context, so the finished clip arrives with its own stereo soundtrack — dialogue, effects, and score shaped in a single forward pass. Expect roughly 15 seconds of footage at up to 2K and 24fps, with every knob exposed at the node level.

  • Audio Generated In-Step
    Dialogue, effects, and music are modeled alongside the picture, so a single pass through the comfyui minimax h3 graph yields one MP4 with everything already in sync.
  • Open Weights, Local Runs
    Because the comfyui minimax h3 model ships as open weights, you can run it on your own hardware and tune duration, resolution, and diffusion settings without hitting an API ceiling.
  • Mix Every Reference Type
    Feed text, pictures, footage, and voice samples into the same job — the comfyui minimax h3 nodes hold a character, look, movement, camera angle, or timbre steady across the take.

Three Steps to Your First comfyui minimax h3 Render

Three quick steps take you from a fresh ComfyUI install to open-weight footage with built-in sound via comfyui minimax h3.

Core Capabilities of the comfyui minimax h3 Workflow

Ship-ready templates, omni-modal reasoning, in-pass stereo audio, reference locking, and an optional Sage Attention boost — together the comfyui minimax h3 workflow covers the whole local video pipeline.

Three Ready-Made Templates

Text-to-video, image-to-video, and reference-to-video ship as separate examples in the comfyui minimax h3 library, so each generation mode works the moment you open it.

One Shared Context Window

Rather than bolting modalities together, the comfyui minimax h3 model reads prompts, pictures, footage, and sound in a single context and blends every reference into one render.

Lock Down What Matters

Hold a face, an art style, a motion path, a camera move, or a voice steady using reference media — the comfyui minimax h3 R2V node accepts as many as 9 images, 3 videos, and 3 audio clips.

Clean On-Screen Text and Logos

Lettering and brand marks come out crisp thanks to the comfyui minimax h3 model, whose instruction following lets you spell out how each reference should relate in plain language.

Faster Runs with Sage Attention

Slot the Patch Sage Attention KJ node into your comfyui minimax h3 graph and generation time roughly halves, with barely any visible drop in quality.

Flexible Resolution and Length Grid

The comfyui minimax h3 Resolution Selector derives width and height from your aspect ratio and megapixel target, snapped to a 32-multiple grid and 17-frame blocks at 24fps.

FAQ

comfyui minimax h3 — Frequently Asked Questions

Answers to the questions people ask most about running MiniMax H3 as a comfyui minimax h3 workflow.

1

What exactly is the comfyui minimax h3 workflow?

It is the native ComfyUI implementation of MiniMax H3, an omni-modal generation model that MiniMax released with open weights. A single forward pass reads your text, images, video, and audio inputs and returns a clip with its own stereo soundtrack.

2

How high can the output resolution go?

Expect up to 2K at 24fps across roughly 15 seconds of footage. The native canvas keeps a 768px short edge, tops out at 768x1344, and rounds dimensions to multiples of 32.

3

Which generation modes ship with it?

Three examples come bundled: text-to-video (T2V), image-to-video (I2V) with optional first and last frame control, and reference-to-video (R2V) for pinning a character, style, motion, camera angle, or voice.

4

Does the model create audio as well?

It does. Voice, sound effects, and music are all modeled natively in stereo alongside the picture, then delivered together inside one synchronized MP4.

5

How do I get started?

Move ComfyUI to 0.30.0 or higher, browse Template Library > Video, select a comfyui minimax h3 example, and let the pop-up pull the weights from the Comfy-Org/MiniMax-H3 repository on Hugging Face.

6

Is there a way to make generation faster?

Install SageAttention along with the KJNodes custom nodes, then insert a Patch Sage Attention KJ node between the UNETLoader and BasicGuider in your comfyui minimax h3 graph — generation time drops by roughly half.

Start Rendering with the comfyui minimax h3 Workflow

Keep everything on your own machine: open weights, stereo sound baked in, and every parameter exposed. Text-to-video, image-to-video, and reference-to-video graphs are ready to load.