MiniMax H3 Open Weights: ComfyUI Setup and the License Catch
TL;DR: MiniMax released H3’s video+audio generation weights in August 2026, and the Comfy-Org repack runs on a 24GB GPU (12GB with heavy offloading). The catch is the license: the H3 Community License explicitly excludes the United States, European Union, United Kingdom, and South Korea — the regions where most self-hosters live. If you’re in one of those four, Wan 2.2 (Apache 2.0) or LTX-2 are the models to run instead.
| MiniMax H3 (local) | Hailuo API (hosted) | Wan 2.2 (Apache 2.0) | |
|---|---|---|---|
| Best for | Joint video+audio generation, if you’re in a licensed territory | US/EU users who just want H3 output | US/EU self-hosters who need a clean license |
| Hardware / cost | 24GB GPU recommended, ~42.5GB download | Pay per generation, no GPU | 5B variant runs on 8GB VRAM |
| The catch | License excludes US, EU, UK, South Korea | Nothing self-hosted, prompts leave your machine | No native audio track |
Honest take: H3 is the most capable open-weight audio+video model released so far, and most readers of this site cannot legally run it. Check the license before you download 42GB — and if you’re in the US or EU, set up Wan 2.2 instead and skip the gray zone entirely.
What did MiniMax actually release in the H3 open-weights drop?
MiniMax published the H3-Base model checkpoints — the FL2VA (first/last-frame-to-video-audio) and Ref2VA (reference-to-video-audio) variants — on Hugging Face at MiniMaxAI/MiniMax-H3 in early August 2026, with native ComfyUI support landing on August 3. The repository includes the processor, tokenizer, text encoder, visual VAE, audio VAE, and the Omni Transformer itself, with documented deployment paths for ComfyUI, diffusers, vLLM, and SGLang.
Architecturally, H3 is not a language model. It’s a single-stream packed-token diffusion transformer that denoises video and stereo audio latents jointly, conditioned on Qwen3-VL-32B hidden states with per-token modality tags. That joint denoising is the headline feature: sound is generated in the same diffusion pass as the picture, not bolted on by a separate audio model afterward.
Two pieces did not ship, and they matter:
- Context-IR, the hosted preprocessing system that translates free-form prompts into the representation H3-Base consumes, stays on MiniMax’s servers. The ComfyUI workflows work without it, but complex multi-shot prompting is weaker locally.
- H3-Regenerate-2K, the upscaling pass, is also hosted-only. Local generation is native 768px on the short edge. If you’ve seen 2K H3 demos, that resolution came from the hosted pipeline, not the open weights.
So “H3 is open source” needs qualifying: the generator is downloadable, the orchestration and upscaling layers are not, and the license is not an open-source license at all — which brings us to the real problem.
Who can legally run MiniMax H3? Not the US, EU, UK, or South Korea
The MiniMax H3 Community License defines its “Applicable Territory” as worldwide excluding the European Union, the United Kingdom, the Republic of Korea, and the United States. Using, modifying, distributing, or hosting the weights — or their outputs — in those four regions is not licensed without separate written authorization from MiniMax (TechTimes, Aug 4 2026).
For a site whose readers are roughly 65% US and 20% EU, that’s not a footnote. It’s the verdict.
Inside the licensed territory, the terms are Llama-style community terms rather than genuine FOSS:
- Commercial use is allowed below $20 million in annual revenue. Above that, you need a separate commercial agreement.
- Attribution is mandatory: a commercial product or service using H3 must display “MiniMax H3” prominently in its UI.
- No distillation, anywhere: the license prohibits using H3 or its outputs to improve any other AI model. This clause is global, not territory-scoped.
This territory-exclusion pattern isn’t unique to MiniMax — Tencent’s Hunyuan video models have shipped with EU/UK/South Korea carve-outs since 2024 — but H3 is the highest-profile release to add the United States to the excluded list. If the license terms on open-weight models affect what you can build, our breakdown of open-weight licenses that block commercial use covers the wider landscape.
One practical wrinkle: the restriction follows where the model runs and where outputs are used, not your passport. Renting a cloud GPU physically located in Singapore doesn’t obviously cure anything if you, the user, are deploying from Berlin — and nothing in the license text suggests it does. We’re not lawyers; the safe reading is that if you’re in one of the four excluded regions, H3 local deployment isn’t licensed for you, full stop.
What hardware does MiniMax H3 need in ComfyUI?
A 24GB GPU is the comfortable tier; 12GB is the realistic floor. The Comfy-Org repack — the officially maintained consumer-GPU route — prunes roughly 40% of the model’s modulation weights into a lookup table and applies INT8 quantization. The download breaks down as:
| Component | Size |
|---|---|
| Pruned INT8 diffusion model | 21 GB |
| NVFP4 text encoder | 15.7 GB |
| Video VAE | 5.21 GB |
| Audio VAE | 605 MB |
| Total | ~42.5 GB |
VRAM tiers, per community testing compiled at kingy.ai’s ComfyUI guide and the minimaxh3.app VRAM writeups (September 2026):
- 24GB (used RTX 3090, $1,150–$1,350 as of September 2026): pruned INT8 build sits fully in VRAM. This is the recommended entry point.
- 12–16GB: workable with the pruned INT8 model plus the NVFP4 text encoder and aggressive offloading — but you need 32GB+ of system RAM, and generation times stretch.
- 8GB: experimental only, via GGUF or INT4 community builds at small canvases. Don’t plan a workflow around it.
System RAM is the requirement people underestimate. Community measurements put peak system RAM at 93.1 GiB for the full INT8 model, 77.6 GiB for pruned INT8, and 71.2 GiB for pruned FP8 during offloaded runs. With DDR5 64GB kits at $680–$1,070 (September 2026 street prices — the DRAM crisis is real), “just add RAM” is now a bigger line item than a used mid-range GPU. A 64GB machine handles the pruned builds; 16GB machines technically run but push a 5-second clip toward 25 minutes.
Speed expectations on the recommended path: roughly 6–10 minutes per 5-second clip at 0.4 MP with the pruned INT8 model and Turbo LoRA. This is not real-time anything. It’s a render farm workload on one card.
On FP8 vs INT8: both pruned builds are 21GB, but the official ComfyUI templates are tuned for INT8, and the one published head-to-head found pruned FP8 slower at every megapixel count. Use INT8 unless you have a specific reason not to.
How do you set up MiniMax H3 in ComfyUI?
H3 support is native in current ComfyUI — no custom node pack required. The setup, assuming you’re in a licensed territory:
# Update ComfyUI to latest (H3 nodes shipped Aug 3 2026)
cd ComfyUI && git pull
pip install -r requirements.txt
# Model files from the Comfy-Org repack go to:
# models/diffusion_models/ <- pruned INT8 diffusion model (21GB)
# models/text_encoders/ <- NVFP4 text encoder (15.7GB)
# models/vae/ <- video VAE (5.21GB) + audio VAE (605MB)
python main.py
# Then: Workflow -> Browse Templates -> Video -> MiniMax H3
Expected first-run behavior: the template loads four H3-specific nodes (model loader, text encoder, dual VAE decode, audio mux), and the first generation takes noticeably longer than subsequent ones while weights page in. If ComfyUI is new to you, start with our ComfyUI review and the custom nodes guide — though again, H3 itself needs no custom nodes, just a current build.
A problem you will actually hit: out-of-memory on the text encoder, not the diffusion model. The NVFP4 text encoder is 15.7GB on its own, and on a 24GB card ComfyUI has to load it, encode, then swap to the diffusion model. If you see OOM during prompt encoding, enable weight offloading for the text encoder (it runs once per prompt, so the speed penalty is small) rather than downgrading the diffusion model — quality lives in the diffusion weights.
Can you run MiniMax H3 with Ollama, vLLM, or Open WebUI?
Not with Ollama or Open WebUI — H3 is a diffusion video model, not an LLM, so the Ollama/llama.cpp GGUF text stack does not apply. There is no ollama pull minimax-h3 and there never will be; community GGUF builds of H3 target ComfyUI’s GGUF loader, which is a different pipeline than llama.cpp despite the shared file format.
vLLM and SGLang do serve H3, but look at the recipe targets before getting excited: the official vLLM recipe lists GB200 NVL4, B300, MI300X-class hardware, with RTX 4090/5090 at the consumer end running reduced configurations. The vLLM path is for API-style batch serving of the full-precision model — datacenter territory. For a home lab, ComfyUI with the pruned INT8 repack is the only sensible route.
If you came here because you conflated H3 with MiniMax’s text models: the LLM side is covered in our MiniMax M3 open-weight review, and that one doesn’t carry the territory exclusion.
What should US and EU self-hosters run instead?
Wan 2.2 is the default answer: Apache 2.0, no territory games, and its 5B variant runs on 8GB VRAM — a lower bar than H3’s 12GB floor. The 14B variant is the quality pick on a 24GB card. It’s natively supported in ComfyUI with official templates, same as H3. What you give up is the joint audio track: Wan generates silent video.
If synchronized audio is the reason H3 caught your eye, LTX-2 generates audio and video in a single diffusion pass and is the closest licensed-everywhere equivalent — verify the exact license terms of the current LTX release before commercial use, as the LTX family has mixed weights/usage terms across versions. Tencent’s HunyuanVideo remains strong on pure visual quality but carries its own EU/UK/South Korea exclusion, so for EU readers it solves nothing.
What to actually buy
Prices as of September 2026, all verified against current street prices:
| Your situation | The move | Price | Where |
|---|---|---|---|
| In a licensed territory, want H3 comfortably | Used RTX 3090 24GB | ~$1,150–$1,350 | Check price |
| In the US/EU — run Wan 2.2 5B instead | Any 8–12GB card you already own | $0 | — |
| Want to test the workload before buying VRAM | Rented RTX 5090, from $0.25/hr | pay per hour | Vast.ai |
On the rental row: Vast.ai lets you filter machines by country, which matters more than usual here — for H3, pick a host outside the four excluded regions, and understand that this addresses where the model runs, not necessarily your own position as the deployer. For Wan or LTX testing, rent anywhere.
Full GPU-buying context, including why 24GB used cards beat 16GB new cards for generative workloads, is in runaihome’s GPU buying guide for local AI.
When not to bother with H3 locally
- You’re in the US, EU, UK, or South Korea. The license doesn’t cover you. Use the hosted Hailuo API if you need H3 output specifically, or self-host Wan 2.2 / LTX-2 with a clean conscience.
- You have under 32GB of system RAM. The offloading path eats 70–93 GiB at peak for the bigger builds; a 16GB machine turns 8-minute renders into 25-minute ones. With DDR5 prices where they are, that upgrade may cost more than the GPU.
- You need 2K output. Local generation tops out at 768px short edge. The 2K pass is hosted-only, so a “fully local 2K H3 pipeline” does not exist regardless of your hardware.
- You expected an LLM. H3 is a video model. For MiniMax’s text models, see the M3 review linked above.
FAQ
Is MiniMax H3 open source? No, by any OSI definition. The weights are downloadable, but the H3 Community License restricts territory (excluding the US, EU, UK, and South Korea), caps commercial use at $20M annual revenue, mandates UI attribution, and bans using outputs to train other models. “Open weights, restrictive license” is the accurate description.
How much VRAM does MiniMax H3 need? 24GB runs the Comfy-Org pruned INT8 build fully in VRAM and is the recommended tier. 12–16GB works with offloading plus 32GB+ system RAM. 8GB is experimental GGUF/INT4 territory at small resolutions. The total model download is about 42.5GB.
Does MiniMax H3 generate audio locally? Yes — this is its differentiator. The open FL2VA and Ref2VA checkpoints denoise video and stereo audio latents jointly in one diffusion pass, and the audio VAE (605MB) ships with the weights. Local clips come out with synchronized sound at native 768px.
Sources
- TechTimes — MiniMax H3 Open Weights Exclude US, EU, UK, and Korea From Local Deployment (Aug 4, 2026)
- MiniMax — Open General Intelligence: MiniMax H3 Is Now Open Source (official announcement)
- MiniMax H3 LICENSE file on Hugging Face
- vLLM official recipe for MiniMax-H3
- kingy.ai — MiniMax H3 ComfyUI Guide: Setup, VRAM, Workflows & Fixes (Sept 2026)
Was this article helpful?
Thanks for the feedback — it helps improve future articles.
Need hands-on help?
I offer 1-on-1 technical consulting for local AI setup, GPU selection, and AI coding tool configuration — same topics covered on this site.
Book a session — $49 / hour →What self-hosting actually costs
Real cost breakdowns for self-hosted AI: hardware floors, power, maintenance hours, and the honest comparison against paying for it. No spam, unsubscribe anytime.