← All guides

Can I Run MiniMax H3 Locally? Yes, 16GB Works (2026)

Hugging Face lists MiniMax H3 at 33.1B parameters. That count is the video transformer only. The model also needs a Qwen3-VL-32B text encoder, and 13B of the transformer's parameters are AdaLN branches you can drop. We read every file size from the Hugging Face API on 25 September 2026, checked ComfyUI, SGLang, DiffSynth and the GGUF repos, and found which part really sets the VRAM floor.

Bottom Line

  • Yes, a 16GB card runs MiniMax H3. One RTX 5070 Ti 16GB user ran ComfyUI’s official template at 768p. A 5-second clip took about 3-5 minutes with 64GB of RAM (community report).
  • 24GB is the comfortable tier. It holds ComfyUI’s pruned INT8 video transformer (20.97 GB) without streaming. SGLang measured a 24GB RTX 4090 D at 164-406 s per 1344x768 clip.
  • The text encoder is bigger than the video model. After pruning, the Qwen3-VL-32B encoder is the largest file in every matched format. On 16GB, it sets the floor.
  • System RAM decides speed. SGLang writes that on consumer hardware “the binding question is not which card you have but how much host RAM sits behind it.”
  • License: the community license excludes the US, EU, UK and South Korea. People there must apply to MiniMax for a license.
  • Ollama does not run it. Use ComfyUI (native since v0.30.0), SGLang, DiffSynth-Studio, diffusers or stable-diffusion.cpp.

Minimum VRAM by path, from the sources below:

PathFiles total (arithmetic)Minimum cardEvidence
Official BF16, unpruned (per task)144.02 GBMulti-GPU, or 32GB with heavy streamingMiniMax’s example uses 4 GPUs; SGLang ran one RTX 5090 with streaming
ComfyUI template (pruned INT8 + NVFP4 encoder)40.08 GB24GB resident, 16GB with offloadComfy docs; 5070 Ti 16GB report
Pruned Q4 GGUF + Q4_K_M encoder31.8-35.6 GB16GBAbiray card recommends Q4_K_M for 16GB
Pruned Q3 GGUF, encoder on CPU23.1-27.8 GB12GBunsloth te=cpu recipe; SGLang 12 GiB recipe
DiffSynth NF4, disk offload27.69 GB8GB (7GB in the docs)DiffSynth repo and docs

What the Model Is

Figures come from the MiniMax-H3 model card and the Hugging Face API, read on 25 September 2026.

SpecValue
Video transformer33B dense, single-stream. About 13B sit in AdaLN branches
Text encoderFull Qwen3-VL-32B. H3 reads the hidden states of layer 50
CheckpointsFL2VA (text, first and last frame) and Ref2VA (up to 9 images, 3 videos, 3 audio clips)
Output4-15 s, 24 fps, 32 kHz stereo audio, 768p short edge by default
2K outputNeeds H3-Regenerate-2K, which is not open-sourced. API only
Prompt refinerH3-Context-IR is hosted only. MiniMax calls it “critical to the quality”
PrecisionBF16, CFG-distilled
HF repo created28 July 2026. 3,611,503 downloads and 5,676 likes on 25 September

The 33,122,992,896 parameters on the Hugging Face page are the transformer only. We summed the safetensors headers: 13.006B of them are adaln_proj weights, one per block for 50 blocks. The model card says the AdaLN outputs “can be precomputed and cached”, so inference does not need these weights. That is where the “pruned” files come from.

Every File Size That Matters

Sizes come from the Hugging Face API with ?blobs=true, read on 25 September 2026. GB means 10^9 bytes.

Video transformer

FileSourceSize (GB)
BF16, fullMiniMaxAI, Comfy-Org66.28
BF16, prunedComfy-Org40.23
INT8 ConvRot, fullComfy-Org34.04
INT8 ConvRot, pruned (ComfyUI template)Comfy-Org20.97
FP8 scaled, prunedComfy-Org20.96
NVFP4, prunedAbiray12.53
NF4, prunedDiffSynth-Studio10.48
Q8_0 GGUF, prunedunsloth / Abiray21.44 / 21.58
Q6_K GGUF, prunedunsloth / Abiray16.59 / 16.73
Q5 GGUF, prunedunsloth Q5_0 / Abiray Q5_K_M13.92 / 14.07
Q4 GGUF, prunedunsloth Q4_K / Abiray Q4_K_M11.42 / 11.56
Q3 GGUF, prunedunsloth Q3_K / Abiray Q3_K_M8.76 / 8.90
Q2 GGUF, prunedunsloth Q2_K / UD-Q2_K_XL6.72 / 8.06
Q4_K_M GGUF, fullAbiray19.86

The full minus the pruned BF16 file is 26.05 GB (arithmetic). That matches 13.0B parameters at 2 bytes each. A full Q4_K_M GGUF (19.86 GB) is almost as large as a pruned Q8_0 (21.44 GB). Always pick “pruned”.

Text encoder (Qwen3-VL-32B, H3 version)

FileSourceSize (GB)
BF16, all 64 layersMiniMaxAI66.71
BF16, cut at layer 50Comfy-Org51.51
INT8 ConvRotComfy-Org27.14
NVFP4 AWQ (ComfyUI template)Comfy-Org15.69
NF4DiffSynth-Studio15.32
Q4_K_M GGUFDeepBeepMeep and Abiray / unsloth14.58 / 18.22
Q2_K GGUFDeepBeepMeep / unsloth Q2_K_M8.49 / 13.10

The official encoder ships all 64 layers, but H3 reads layer 50. Comfy’s copy is 15.20 GB smaller (arithmetic). Our count of the 14 unused layers plus the output head is 7.6B parameters, or 15.2 GB at BF16. The numbers agree.

Comfy-Org says its NVFP4 encoder “does not require Blackwell GPU to use.” Stock Qwen3-VL-32B chat GGUFs are a risk: the model card says H3 adds special tokens and needs its own tokenizer files.

VAEs

FileSize (GB)
Video VAE, official10.42
Video VAE, FP16 (Comfy-Org)5.21
Video VAE, INT8 ConvRot (ComfyUI template)2.81
Video VAE, NF4 (DiffSynth)1.61
Audio VAE, FP320.61
Turbo LoRA, ComfyUI (optional)1.96

The video VAE decoder is a 36-layer ViT. That is why it is 10.42 GB, much larger than most video VAEs.

The fact most VRAM guides leave out

After pruning, the text encoder is larger than the video model at every matched precision:

PrecisionPruned transformerEncoder
BF1640.23 GB51.51 GB
INT8 ConvRot20.97 GB27.14 GB
NF4 (DiffSynth)10.48 GB15.32 GB
Q4_K_M GGUF (Abiray)11.56 GB14.58 GB

ComfyUI encodes the prompt first, then samples. So each phase needs its own part in VRAM, not both at once. On 16GB, the Q4 video model fits with room left. The Q4 encoder (14.58 GB) or the NVFP4 encoder (15.69 GB) almost fills the card by itself. So the encoder sets the floor, not the “33B” video model. unsloth’s stable-diffusion.cpp recipe solves this with --backend te=cpu, “which keeps the 12 GB text encoder off the card.”

Decoder Table: Your GPU, Your Files

Method: file sizes above, plus the encode-then-sample order. Activations add more, and they grow with resolution and clip length. Rows marked “measured” name a source.

VRAMVideo modelText encoderExpected experience
8GBNF4 pruned (10.48 GB), streamedNF4 (15.32 GB), streamedDiffSynth says “as little as 8GB”, with weights offloaded to disk. Its docs list 7GB as the minimum. Slow. We found no timed 8GB run.
12GBQ3 pruned GGUF (8.76-8.90 GB)Q2 GGUF on CPUMeasured (SGLang): a 12 GiB cap with 32GB host RAM ran 864x480, 124 frames, 20 steps in 319-356 s. The cap was simulated in a lab.
16GBQ4 pruned GGUF (11.42-11.56 GB), or template INT8 (20.97 GB) with offloadNVFP4 (15.69 GB) or Q4_K_M GGUF (14.58 GB)Measured (community): RTX 5070 Ti, official template, about 3-5 min per 5 s 768p clip. 32GB of RAM froze during load; 64GB fixed it.
24GBINT8 pruned (20.97 GB) or Q8_0 pruned (21.44 GB), residentNVFP4 or INT8 (27.14 GB) with offloadMeasured (SGLang): RTX 4090 D, 1344x768, 107 frames (about 4.5 s), 20 steps, 164-406 s depending on attention backend. GPU peak about 18 GB.
32GBFull BF16 streamed, or INT8 pruned residentAnyMeasured (SGLang): RTX 5090, 60GB RAM, 864x480, 124 frames, 112 s. ComfyUI took 141-146 s on the same box.
96GBBF16 pruned (40.23 GB) residentBF16 layer-50 (51.51 GB), swappedBoth at BF16 total 91.74 GB (arithmetic). Fits on paper, with little room for activations. We found no measured run.
128GB unified (DGX Spark)Automatic offloadAutomaticMeasured (SGLang): about 12 min of load, then about 12 min per warm 480p request.

Why 24GB is not the floor

The ComfyUI template’s video model is 20.97 GB, so many guides call 24GB the minimum. The RTX 5070 Ti report ran the same files on 16GB. ComfyUI streamed the part that did not fit. The cost was time and system RAM, not a crash.

Why RAM matters as much as VRAM

SGLang’s 16 GiB recipe (Recipe B) was fast: 120.92 s per 864x480 clip. But it pinned 116.7 GB of host memory. On the RTX 5090 desktop with 60GB of RAM, SGLang reports that ComfyUI’s default command was “OOM-killed during load” and needed --fast-disk. That run used the unpruned BF16 files. Plan on 64GB of RAM for 16GB and 24GB cards, and more for full BF16 files.

Which Tools Run It

We checked each runtime on 25 September 2026.

RuntimeStatus
ComfyUINative. “Support MiniMax-H3” (PR #15224) shipped in v0.30.0 on 3 August 2026. The Comfy tutorial templates cover T2V, I2V and R2V. Default is 20 steps; turbo mode uses an 8-step LoRA.
ComfyUI + GGUFAbiray’s pruned GGUFs load with the ComfyUI-GGUF UnetLoaderGGUF node.
SGLangYes. MiniMax’s own example. The SGLang cookbook has consumer recipes for 12GB to 32GB cards.
diffusers, vLLMYes, listed on the model card.
DiffSynth-StudioYes, with NF4 files and automatic VRAM management.
stable-diffusion.cppYes, per unsloth. --cfg-scale 1.0 is required; H3 aborts above 1.0.
OllamaNo. ollama.com/library/minimax-h3 returns 404. Ollama’s MiniMax entries are M-series text models.

Comfy’s docs say Sage Attention can “roughly double the generation speed with minimal quality loss.” SGLang’s RTX 4090 D table shows the trade: Sage cut time by 2.32x, and pixel fidelity to BF16 fell to 23.51 dB PSNR.

Download only what you need

Do not download the whole MiniMax repo. It holds the original and diffusers formats side by side: 498.47 GB of files (arithmetic). The model card shows hf download ... --include "FL2VA/*" for one task family. For ComfyUI, the four template files total 40.08 GB (arithmetic).

Which GPU to Buy for It

Prices come from our tracked price file, as of September 2026.

Read this first if you live in the US, EU, UK or South Korea. The license does not cover you unless MiniMax approves your use. Do not buy a card for H3 alone. The same 24GB and 32GB cards also run LTX-2.5 and Qwen-Image 2.1.

  • Best fit: 24GB. A used RTX 3090 24GB (about $1,463 average used) holds the 20.97 GB pruned INT8 video model without streaming. It is the same VRAM class as the RTX 4090 D in SGLang’s test, but slower. Pair it with 64GB of RAM.
  • Faster and longer clips: 32GB. The RTX 5090 32GB (from $4,299) ran a 480p clip in 112 s in SGLang’s desktop test.
  • Budget: 16GB. The RTX 5060 Ti 16GB ($789 at B&H) is in the same VRAM class as the RTX 5070 Ti report. We found no 5060 Ti run, and it has less compute. Expect several minutes per clip.
  • Rent first. RunPod lists an RTX 4090 at $0.74/hr Secure or $0.34/hr Community, and an RTX 5090 at $0.99/$0.69. Test a clip there before you buy. Check the license territory first.

AMD and Mac: SGLang lists AMD Instinct data-center runs only. We found no measured Radeon or Apple Silicon run. For other tiers, see the tested hardware list.

Honest Caveats

  1. We did not run the model. Every speed number is from SGLang’s docs or one named community report.
  2. Most SGLang consumer numbers used a simulated VRAM cap. SGLang says so. The RTX 5090 desktop run and the RTX 4090 D run were real cards.
  3. The local model is not the full product. Open weights give 768p H3-Base. 2K output and the Context-IR prompt refiner are API only.
  4. The license limits where you can use it. The US, EU, UK and South Korea are Excluded Territories. The license also needs written approval above $20M yearly revenue. Comfy’s docs add that commercial use of local outputs needs a MiniMax commercial license, sold through Comfy. Read the license and its Q&A.
  5. The decoder table is arithmetic. Activations for 15 s clips add memory we cannot pin down.
  6. Sparse attention is not released yet. MiniMax says the open release uses full attention only. Speeds may change when sparse attention ships.

Sources

See Also

Need OpenClaw fixed live?

Remote rescue sessions for gateway, auth, tunnel, VPS, and model access problems.

See Rescue Session

Read next

Best Qwen Model for 16GB VRAM (2026): Qwen3.8 27B Wins
Best Qwen model for 16GB VRAM: Qwen3.8-27B at 3.5 bits (10.96 GiB). The Q4_K_M 27B files do not fit. Byte-exact GGUF sizes, KV math, and the best Qwen for 8, 12, 24 and 32GB.
Can I Run LTX-2.5 Locally? VRAM by GPU (2026)
LTX-2.5 is a 22B video model plus a 12B Gemma 4 text encoder. The official INT8 files need a 24GB card. 16GB works with NVFP4 or GGUF and weight streaming. The '16GB minimum' leans on cloud text encoding. Ollama does not run it.
Is a Used RTX 3090 Worth $1,450 for Local LLMs? (2026)
A used RTX 3090 sells for a $1,463 average as of September 2026, about 97% of its $1,499 launch price in 2020. It is no longer the cheapest VRAM per GB, but it is still the cheapest memory bandwidth. A verdict per buyer, what $1,450 buys instead, and the rent-vs-buy hours.
Can I Run Qwen-Image 2.1 Locally? Yes, on 12GB
Qwen-Image 2.1 is a 7B image model, but its Qwen3-VL 8B text encoder is bigger than the image model. A Q4 GGUF stack is 9.91 GB, so 12GB works. 16GB runs INT8 at about 20 s per image. Ollama does not run it.