Can I Run MiniMax H3 Locally? Yes, 16GB Works (2026)
Hugging Face lists MiniMax H3 at 33.1B parameters. That count is the video transformer only. The model also needs a Qwen3-VL-32B text encoder, and 13B of the transformer's parameters are AdaLN branches you can drop. We read every file size from the Hugging Face API on 25 September 2026, checked ComfyUI, SGLang, DiffSynth and the GGUF repos, and found which part really sets the VRAM floor.
Bottom Line
- Yes, a 16GB card runs MiniMax H3. One RTX 5070 Ti 16GB user ran ComfyUI’s official template at 768p. A 5-second clip took about 3-5 minutes with 64GB of RAM (community report).
- 24GB is the comfortable tier. It holds ComfyUI’s pruned INT8 video transformer (20.97 GB) without streaming. SGLang measured a 24GB RTX 4090 D at 164-406 s per 1344x768 clip.
- The text encoder is bigger than the video model. After pruning, the Qwen3-VL-32B encoder is the largest file in every matched format. On 16GB, it sets the floor.
- System RAM decides speed. SGLang writes that on consumer hardware “the binding question is not which card you have but how much host RAM sits behind it.”
- License: the community license excludes the US, EU, UK and South Korea. People there must apply to MiniMax for a license.
- Ollama does not run it. Use ComfyUI (native since v0.30.0), SGLang, DiffSynth-Studio, diffusers or stable-diffusion.cpp.
Minimum VRAM by path, from the sources below:
| Path | Files total (arithmetic) | Minimum card | Evidence |
|---|---|---|---|
| Official BF16, unpruned (per task) | 144.02 GB | Multi-GPU, or 32GB with heavy streaming | MiniMax’s example uses 4 GPUs; SGLang ran one RTX 5090 with streaming |
| ComfyUI template (pruned INT8 + NVFP4 encoder) | 40.08 GB | 24GB resident, 16GB with offload | Comfy docs; 5070 Ti 16GB report |
| Pruned Q4 GGUF + Q4_K_M encoder | 31.8-35.6 GB | 16GB | Abiray card recommends Q4_K_M for 16GB |
| Pruned Q3 GGUF, encoder on CPU | 23.1-27.8 GB | 12GB | unsloth te=cpu recipe; SGLang 12 GiB recipe |
| DiffSynth NF4, disk offload | 27.69 GB | 8GB (7GB in the docs) | DiffSynth repo and docs |
What the Model Is
Figures come from the MiniMax-H3 model card and the Hugging Face API, read on 25 September 2026.
| Spec | Value |
|---|---|
| Video transformer | 33B dense, single-stream. About 13B sit in AdaLN branches |
| Text encoder | Full Qwen3-VL-32B. H3 reads the hidden states of layer 50 |
| Checkpoints | FL2VA (text, first and last frame) and Ref2VA (up to 9 images, 3 videos, 3 audio clips) |
| Output | 4-15 s, 24 fps, 32 kHz stereo audio, 768p short edge by default |
| 2K output | Needs H3-Regenerate-2K, which is not open-sourced. API only |
| Prompt refiner | H3-Context-IR is hosted only. MiniMax calls it “critical to the quality” |
| Precision | BF16, CFG-distilled |
| HF repo created | 28 July 2026. 3,611,503 downloads and 5,676 likes on 25 September |
The 33,122,992,896 parameters on the Hugging Face page are the transformer only. We summed the safetensors headers: 13.006B of them are adaln_proj weights, one per block for 50 blocks. The model card says the AdaLN outputs “can be precomputed and cached”, so inference does not need these weights. That is where the “pruned” files come from.
Every File Size That Matters
Sizes come from the Hugging Face API with ?blobs=true, read on 25 September 2026. GB means 10^9 bytes.
Video transformer
| File | Source | Size (GB) |
|---|---|---|
| BF16, full | MiniMaxAI, Comfy-Org | 66.28 |
| BF16, pruned | Comfy-Org | 40.23 |
| INT8 ConvRot, full | Comfy-Org | 34.04 |
| INT8 ConvRot, pruned (ComfyUI template) | Comfy-Org | 20.97 |
| FP8 scaled, pruned | Comfy-Org | 20.96 |
| NVFP4, pruned | Abiray | 12.53 |
| NF4, pruned | DiffSynth-Studio | 10.48 |
| Q8_0 GGUF, pruned | unsloth / Abiray | 21.44 / 21.58 |
| Q6_K GGUF, pruned | unsloth / Abiray | 16.59 / 16.73 |
| Q5 GGUF, pruned | unsloth Q5_0 / Abiray Q5_K_M | 13.92 / 14.07 |
| Q4 GGUF, pruned | unsloth Q4_K / Abiray Q4_K_M | 11.42 / 11.56 |
| Q3 GGUF, pruned | unsloth Q3_K / Abiray Q3_K_M | 8.76 / 8.90 |
| Q2 GGUF, pruned | unsloth Q2_K / UD-Q2_K_XL | 6.72 / 8.06 |
| Q4_K_M GGUF, full | Abiray | 19.86 |
The full minus the pruned BF16 file is 26.05 GB (arithmetic). That matches 13.0B parameters at 2 bytes each. A full Q4_K_M GGUF (19.86 GB) is almost as large as a pruned Q8_0 (21.44 GB). Always pick “pruned”.
Text encoder (Qwen3-VL-32B, H3 version)
| File | Source | Size (GB) |
|---|---|---|
| BF16, all 64 layers | MiniMaxAI | 66.71 |
| BF16, cut at layer 50 | Comfy-Org | 51.51 |
| INT8 ConvRot | Comfy-Org | 27.14 |
| NVFP4 AWQ (ComfyUI template) | Comfy-Org | 15.69 |
| NF4 | DiffSynth-Studio | 15.32 |
| Q4_K_M GGUF | DeepBeepMeep and Abiray / unsloth | 14.58 / 18.22 |
| Q2_K GGUF | DeepBeepMeep / unsloth Q2_K_M | 8.49 / 13.10 |
The official encoder ships all 64 layers, but H3 reads layer 50. Comfy’s copy is 15.20 GB smaller (arithmetic). Our count of the 14 unused layers plus the output head is 7.6B parameters, or 15.2 GB at BF16. The numbers agree.
Comfy-Org says its NVFP4 encoder “does not require Blackwell GPU to use.” Stock Qwen3-VL-32B chat GGUFs are a risk: the model card says H3 adds special tokens and needs its own tokenizer files.
VAEs
| File | Size (GB) |
|---|---|
| Video VAE, official | 10.42 |
| Video VAE, FP16 (Comfy-Org) | 5.21 |
| Video VAE, INT8 ConvRot (ComfyUI template) | 2.81 |
| Video VAE, NF4 (DiffSynth) | 1.61 |
| Audio VAE, FP32 | 0.61 |
| Turbo LoRA, ComfyUI (optional) | 1.96 |
The video VAE decoder is a 36-layer ViT. That is why it is 10.42 GB, much larger than most video VAEs.
The fact most VRAM guides leave out
After pruning, the text encoder is larger than the video model at every matched precision:
| Precision | Pruned transformer | Encoder |
|---|---|---|
| BF16 | 40.23 GB | 51.51 GB |
| INT8 ConvRot | 20.97 GB | 27.14 GB |
| NF4 (DiffSynth) | 10.48 GB | 15.32 GB |
| Q4_K_M GGUF (Abiray) | 11.56 GB | 14.58 GB |
ComfyUI encodes the prompt first, then samples. So each phase needs its own part in VRAM, not both at once. On 16GB, the Q4 video model fits with room left. The Q4 encoder (14.58 GB) or the NVFP4 encoder (15.69 GB) almost fills the card by itself. So the encoder sets the floor, not the “33B” video model. unsloth’s stable-diffusion.cpp recipe solves this with --backend te=cpu, “which keeps the 12 GB text encoder off the card.”
Decoder Table: Your GPU, Your Files
Method: file sizes above, plus the encode-then-sample order. Activations add more, and they grow with resolution and clip length. Rows marked “measured” name a source.
| VRAM | Video model | Text encoder | Expected experience |
|---|---|---|---|
| 8GB | NF4 pruned (10.48 GB), streamed | NF4 (15.32 GB), streamed | DiffSynth says “as little as 8GB”, with weights offloaded to disk. Its docs list 7GB as the minimum. Slow. We found no timed 8GB run. |
| 12GB | Q3 pruned GGUF (8.76-8.90 GB) | Q2 GGUF on CPU | Measured (SGLang): a 12 GiB cap with 32GB host RAM ran 864x480, 124 frames, 20 steps in 319-356 s. The cap was simulated in a lab. |
| 16GB | Q4 pruned GGUF (11.42-11.56 GB), or template INT8 (20.97 GB) with offload | NVFP4 (15.69 GB) or Q4_K_M GGUF (14.58 GB) | Measured (community): RTX 5070 Ti, official template, about 3-5 min per 5 s 768p clip. 32GB of RAM froze during load; 64GB fixed it. |
| 24GB | INT8 pruned (20.97 GB) or Q8_0 pruned (21.44 GB), resident | NVFP4 or INT8 (27.14 GB) with offload | Measured (SGLang): RTX 4090 D, 1344x768, 107 frames (about 4.5 s), 20 steps, 164-406 s depending on attention backend. GPU peak about 18 GB. |
| 32GB | Full BF16 streamed, or INT8 pruned resident | Any | Measured (SGLang): RTX 5090, 60GB RAM, 864x480, 124 frames, 112 s. ComfyUI took 141-146 s on the same box. |
| 96GB | BF16 pruned (40.23 GB) resident | BF16 layer-50 (51.51 GB), swapped | Both at BF16 total 91.74 GB (arithmetic). Fits on paper, with little room for activations. We found no measured run. |
| 128GB unified (DGX Spark) | Automatic offload | Automatic | Measured (SGLang): about 12 min of load, then about 12 min per warm 480p request. |
Why 24GB is not the floor
The ComfyUI template’s video model is 20.97 GB, so many guides call 24GB the minimum. The RTX 5070 Ti report ran the same files on 16GB. ComfyUI streamed the part that did not fit. The cost was time and system RAM, not a crash.
Why RAM matters as much as VRAM
SGLang’s 16 GiB recipe (Recipe B) was fast: 120.92 s per 864x480 clip. But it pinned 116.7 GB of host memory. On the RTX 5090 desktop with 60GB of RAM, SGLang reports that ComfyUI’s default command was “OOM-killed during load” and needed --fast-disk. That run used the unpruned BF16 files. Plan on 64GB of RAM for 16GB and 24GB cards, and more for full BF16 files.
Which Tools Run It
We checked each runtime on 25 September 2026.
| Runtime | Status |
|---|---|
| ComfyUI | Native. “Support MiniMax-H3” (PR #15224) shipped in v0.30.0 on 3 August 2026. The Comfy tutorial templates cover T2V, I2V and R2V. Default is 20 steps; turbo mode uses an 8-step LoRA. |
| ComfyUI + GGUF | Abiray’s pruned GGUFs load with the ComfyUI-GGUF UnetLoaderGGUF node. |
| SGLang | Yes. MiniMax’s own example. The SGLang cookbook has consumer recipes for 12GB to 32GB cards. |
| diffusers, vLLM | Yes, listed on the model card. |
| DiffSynth-Studio | Yes, with NF4 files and automatic VRAM management. |
| stable-diffusion.cpp | Yes, per unsloth. --cfg-scale 1.0 is required; H3 aborts above 1.0. |
| Ollama | No. ollama.com/library/minimax-h3 returns 404. Ollama’s MiniMax entries are M-series text models. |
Comfy’s docs say Sage Attention can “roughly double the generation speed with minimal quality loss.” SGLang’s RTX 4090 D table shows the trade: Sage cut time by 2.32x, and pixel fidelity to BF16 fell to 23.51 dB PSNR.
Download only what you need
Do not download the whole MiniMax repo. It holds the original and diffusers formats side by side: 498.47 GB of files (arithmetic). The model card shows hf download ... --include "FL2VA/*" for one task family. For ComfyUI, the four template files total 40.08 GB (arithmetic).
Which GPU to Buy for It
Prices come from our tracked price file, as of September 2026.
Read this first if you live in the US, EU, UK or South Korea. The license does not cover you unless MiniMax approves your use. Do not buy a card for H3 alone. The same 24GB and 32GB cards also run LTX-2.5 and Qwen-Image 2.1.
- Best fit: 24GB. A used RTX 3090 24GB (about $1,463 average used) holds the 20.97 GB pruned INT8 video model without streaming. It is the same VRAM class as the RTX 4090 D in SGLang’s test, but slower. Pair it with 64GB of RAM.
- Faster and longer clips: 32GB. The RTX 5090 32GB (from $4,299) ran a 480p clip in 112 s in SGLang’s desktop test.
- Budget: 16GB. The RTX 5060 Ti 16GB ($789 at B&H) is in the same VRAM class as the RTX 5070 Ti report. We found no 5060 Ti run, and it has less compute. Expect several minutes per clip.
- Rent first. RunPod lists an RTX 4090 at $0.74/hr Secure or $0.34/hr Community, and an RTX 5090 at $0.99/$0.69. Test a clip there before you buy. Check the license territory first.
AMD and Mac: SGLang lists AMD Instinct data-center runs only. We found no measured Radeon or Apple Silicon run. For other tiers, see the tested hardware list.
Honest Caveats
- We did not run the model. Every speed number is from SGLang’s docs or one named community report.
- Most SGLang consumer numbers used a simulated VRAM cap. SGLang says so. The RTX 5090 desktop run and the RTX 4090 D run were real cards.
- The local model is not the full product. Open weights give 768p H3-Base. 2K output and the Context-IR prompt refiner are API only.
- The license limits where you can use it. The US, EU, UK and South Korea are Excluded Territories. The license also needs written approval above $20M yearly revenue. Comfy’s docs add that commercial use of local outputs needs a MiniMax commercial license, sold through Comfy. Read the license and its Q&A.
- The decoder table is arithmetic. Activations for 15 s clips add memory we cannot pin down.
- Sparse attention is not released yet. MiniMax says the open release uses full attention only. Speeds may change when sparse attention ships.
Sources
- MiniMaxAI/MiniMax-H3 model card: architecture, 13B AdaLN, layer-50 encoder, open vs hosted modules
- Hugging Face API: MiniMaxAI/MiniMax-H3: file sizes, parameter count, dates, downloads; transformer safetensors headers for the
adaln_projcount - MiniMax H3 Community License and License Q&A: Excluded Territories, $20M threshold
- Comfy-Org/MiniMax-H3: ComfyUI file sizes, NVFP4 encoder note
- ComfyUI MiniMax H3 docs: v0.30.0 requirement, template files, resolution, Sage Attention, commercial-output note
- ComfyUI PR #15224 and v0.30.0 release: native support, 3 August 2026
- SGLang MiniMax-H3 cookbook: RTX 4090 D, RTX 5090, 12/16 GiB recipes, DGX Spark timings
- unsloth/MiniMax-H3-GGUF: GGUF sizes, stable-diffusion.cpp flags, encoder on CPU
- Abiray/MiniMax-H3-Pruned-GGUF, Abiray/MiniMax-H3-GGUF: pruned and full GGUF sizes, 16GB recommendation
- DeepBeepMeep/MiniMax-H3: layer-50 encoder files, Q2_K and Q4_K_M encoder GGUFs
- DiffSynth-Studio/MiniMax-H3-NF4 and DiffSynth docs: NF4 sizes, 8GB and 7GB claims
- hiro, RTX 5070 Ti guide (note.com), 7 August 2026: 16GB run, 32GB to 64GB RAM
- HF forum: MiniMax-H3 quantisation RAM/VRAM, 4 August 2026: over 80GB of RAM used with 8-bit files in diffusers
See Also
- Can I run LTX-2.5 locally?: the other open audio-video model, where the Gemma 4 encoder is 37-43% of the files
- Can I run Qwen-Image 2.1 locally?: the image model where a Qwen3-VL encoder is also the bigger half
- Best local LLM for an RTX 3090: what else the 24GB card runs
- Best local LLM for an RTX 5090: the 32GB tier
- Is 16GB of VRAM still enough in 2026?: the tier-level verdict
- Is 32GB of VRAM enough in 2026?: what the step to 32GB buys
- GGUF quant names explained: what Q4_K_M, Q3_K and Q8_0 mean
Need OpenClaw fixed live?
Remote rescue sessions for gateway, auth, tunnel, VPS, and model access problems.
See Rescue Session