← All guides

Can I Run Qwen-Image 2.1 Locally? Yes, on 12GB

Every Qwen-Image 2.1 GGUF page lists the image model alone: 4.20 GB at Q4_K_M. That number leaves out the bigger half. The model needs a Qwen3-VL 8B text encoder, and at BF16 that encoder is 17.53 GB, larger than the 14.23 GB image model. We read every file size from the Hugging Face API on 23 September 2026, checked ComfyUI, stable-diffusion.cpp and Ollama, and collected the measured VRAM reports with their sources.

Bottom Line

  • Yes, you can run it locally. A 12GB card runs the Q4 GGUF stack. A 16GB card runs the INT8 stack without offload. A 6GB card runs it slowly with CPU offload.
  • The GGUF size is not the VRAM answer. The Q4_K_M GGUF is 4.20 GB, but that file is the image model only. Add the text encoder and VAE and the Q4 stack is 9.91 GB.
  • The text encoder is the biggest part. Qwen3-VL 8B is 17.53 GB at BF16. The 7B image model is 14.23 GB at BF16. Most guides skip this.
  • Measured on 16GB: INT8 image model plus INT8 encoder peaked at about 15 GB on an RTX 4060 Ti 16GB, at about 20 s per image (community report).
  • Ollama does not run it. Use ComfyUI (native since v0.37.0), diffusers, or stable-diffusion.cpp.
  • License: Qwen Research License, non-commercial only.

Minimum VRAM by format, derived from the file sizes below:

Format (image model + text encoder)Files totalMinimum card
Q4_K_M GGUF + Q4_K_M encoder + VAE9.91 GB12GB (6-8GB with offload)
Q8_0 GGUF + Q4_K_M encoder + VAE13.35 GB16GB
INT8 + INT8 encoder + VAE (Comfy-Org)17.29 GB16GB (measured, ~15 GB peak)
8-bit, all weights resident (diffusers)16.6 GB resident24GB (measured 21.3 GB peak)
BF16 + BF16 encoder + VAE32.44 GB32GB+ with offload, or 48GB

What the Model Is

All figures come from the Qwen/Qwen-Image-2.1 model card, its config.json files and the Hugging Face API, read on 23 September 2026.

SpecValue
Image model7B parameters, 32 single-stream DiT layers, 32 heads x 128 dims
Text encoderQwen3-VL 8B (Qwen3VLForConditionalGeneration, 36 layers, hidden size 4096)
VAEAutoencoderKLQwenImage21, its own VAE; earlier Qwen-Image VAEs do not work
TasksText-to-image, image editing with up to 10 reference images, transparent RGBA output
Native resolution2048x2048 (1:1), up to 2752x1536 (16:9)
Default steps40 in the diffusers example
LicenseQwen Research License Agreement, non-commercial
Created on Hugging Face14 September 2026

The model card says “7B parameters in its visual generation component.” The word “component” matters. The text encoder is a separate 8B vision-language model. It reads your prompt and any reference images before the image model starts.

Every File Size That Matters

Sizes come from the Hugging Face API with ?blobs=true, read on 23 September 2026. GB means 10^9 bytes.

Image model (the “diffusion model”)

FileSourceSize (GB)
BF16 safetensorsComfy-Org14.23
INT8 ConvRotComfy-Org7.26
FP8unsloth/Qwen-Image-2.1-FP87.12
INT8unsloth/Qwen-Image-2.1-FP87.26
F16 GGUFunsloth/Qwen-Image-2.1-GGUF14.23
Q8_0 GGUFunsloth7.64
Q6_K GGUFunsloth6.27
Q5_K_M GGUFunsloth5.39
Q4_K_M GGUFunsloth4.20
Q3_K_M GGUFunsloth3.17
Q2_K GGUFunsloth2.47
Q8_0 / Q6_K / Q4_K / Q2_K GGUFleejet (stable-diffusion.cpp author)7.69 / 6.00 / 4.20 / 2.56

Text encoder (Qwen3-VL 8B)

FileSourceSize (GB)
BF16 safetensorsComfy-Org17.53
INT8 ConvRotComfy-Org9.35
W4A8Comfy-Org6.31
FP8unsloth/Qwen-Image-2.1-FP89.39
Q8_0 GGUFunsloth/Qwen3-VL-8B-Instruct-GGUF8.71
Q6_K GGUFunsloth6.73
UD-Q4_K_XL GGUFunsloth (the encoder Unsloth recommends)5.15
Q4_K_M GGUFunsloth5.03
mmproj F16 (vision part, needed for editing with a GGUF encoder)unsloth1.16

VAE

The ComfyUI VAE file is 0.68 GB (qwen_image_2.1_vae_bf16.safetensors). One ComfyUI user reported an error with the VAE linked from the Unsloth card and fixed it with the official file (discussion).

The fact the GGUF pages leave out

At every precision, the text encoder is as big as the image model or bigger. At BF16 it is 17.53 GB against 14.23 GB. At 8 bits it is 9.35 GB against 7.26 GB. So “Q4_K_M = 4.20 GB” describes less than half of what you load. The Unsloth card does say the GGUF “is the denoiser only,” but the file list still leads with 4.20 GB.

There is a second cost. The model reuses a “prefix KV cache” for the prompt and reference images. The stable-diffusion.cpp docs state that a 4,096-token prefix takes about 4 GiB per condition in FP32. Positive and negative prompts keep separate caches. Reference images add prefix tokens, so editing costs more memory than plain text-to-image.

Decoder Table: Your GPU, Your Files

This table is our estimate. We added the file sizes above and checked the result against the measured reports in the next section. Activations at 1024x1024 add about 2-5 GB on top of resident weights, based on the RTX 5090 diffusers numbers (16.6 GB weights, 21.3 GB peak).

VRAMImage modelText encoderExpected experience
6-8GBQ4_0 or Q4_K_M GGUF (4.20 GB)Q4_K_M GGUF (5.03 GB), offloaded to CPUWorks with --offload-to-cpu. Measured on a 6GB RTX 3050: about 4 min 50 s per 1184x1184 image, 25 steps. The tester had 16 GB of system RAM.
12GBQ4_K_M (4.20 GB) or Q6_K (6.27 GB) GGUFQ4_K_M (5.03 GB) or UD-Q4_K_XL (5.15 GB)Q4 stack is 9.91 GB, so it fits at 1024x1024, batch 1. Unsloth lists this tier. No published speed on a 12GB card yet.
16GBINT8 ConvRot (7.26 GB) or Q8_0 GGUF (7.64 GB)INT8 ConvRot (9.35 GB) or W4A8 (6.31 GB)Measured on an RTX 4060 Ti 16GB with INT8 files: ~15 GB peak, ~20 s per image, ~1 min per multi-image edit.
24GBINT8/FP8 (7.1-7.3 GB) or BF16 (14.23 GB)INT8/FP8 (9.4 GB)8-bit with everything resident peaks at 21.3 GB (1024x1024) and 22.6 GB (2048x2048). A two-reference edit peaked at 25.6 GB, over 24GB.
32GBBF16 (14.23 GB)INT8 (9.35 GB) or BF16 with offloadHolds the 8-bit two-reference edit (25.6 GB). The full BF16 stack is 32.44 GB of files, so BF16 everywhere still needs offload.

Why the 16GB row works with 17.29 GB of files. The measured peak was about 15 GB with no offload flag. Our reading: ComfyUI runs the text encoder, then loads the image model, so both never sit in VRAM at full size together. The reporter did not state this; it is our inference from the numbers.

GGUF or INT8 on an NVIDIA card? An Unsloth team member wrote that GGUF “is naturally much slower and only meant for unified memory/CPU devices,” and that ComfyUI defaults to INT4 or INT8 (discussion). On a 16GB+ NVIDIA card, start with the Comfy-Org INT8 files. Use GGUF when VRAM forces it.

Measured Reports, With Sources

We found three measured reports as of 23 September 2026. We did not run these tests ourselves.

GPUSetupResultSource
RTX 5090 32GBdiffusers, fp8 (torchao) image model + encoder, all on GPU, 40 stepsWeights 16.6 GB; peak 21.3 GB at 1024x1024, 22.6 GB at 2048x2048; 19.3 s at 1024x1024, 115.7 s at 2048x2048HF discussion #32, 22 Sep 2026
RTX 5090 32GBdiffusers, NF4 (bitsandbytes), all on GPU, 40 stepsWeights 10.6 GB; peak 15.2 GB at 1024x1024, 15.4 GB at 2048x2048; 19.2 s at 1024x1024same
RTX 4060 Ti 16GBComfyUI, INT8 ConvRot image model + INT8 Qwen3-VL 8B + BF16 VAE, no offloadPeak ~15 GB, typical 13-14 GB; ~20 s per image; ~1 min per multi-image editHF discussion #35, 22 Sep 2026
RTX 3050 6GB, 16 GB RAMstable-diffusion.cpp, Q4_0 GGUF + Q4_K_M encoder, --offload-to-cpu --diffusion-fa, CFG 1, 25 steps, 1184x1184~4 min 50 s per imageleejet discussion #3, 23 Sep 2026

Three details from these reports change how you should run it:

  1. 2048x2048 costs 6x the time of 1024x1024. The RTX 5090 report measured 19.3 s against 115.7 s. Four times the pixels gave six times the time. Start at 1024x1024.
  2. Reference images cost VRAM. Two reference images raised the 8-bit peak from 21.3 GB to 25.6 GB. The 4-bit peak rose from 15.2 GB to 19.1 GB.
  3. CFG 1 halves the work. The RTX 3050 tester found CFG above 1 runs every step twice. CFG 1 was about twice as fast. The ComfyUI template also uses CFG 1. --diffusion-fa cut step time by 35% on that card.

Which Tools Run It

We checked each runtime on 23 September 2026.

RuntimeStatus
ComfyUINative. PR #16400 “Qwen-image 2.1 support” merged 19 Sep 2026, shipped in v0.37.0 (21 Sep 2026). Official text-to-image and edit templates exist.
ComfyUI + GGUFNeeds a custom node. The stock city96 ComfyUI-GGUF node fails with “Unknown model architecture.” leejet recommends leejet/ComfyUI-GGUF; a community add-on also fixes it.
diffusersYes, QwenImage21Pipeline, from diffusers main (git install). Needs transformers>=5.17. enable_model_cpu_offload() is the official low-memory option.
stable-diffusion.cppYes. Official docs page, GGUF files from leejet, --offload-to-cpu for small cards.
Unsloth DesktopYes, GGUF and FP8, text-to-image and editing.
OllamaNo. See below.
llama.cpp, LM StudioNot listed by any model card we read. These are text-model runners. Do not assume support.
Apple SiliconMLX conversions exist (for example ddalcu/Qwen-Image-2.1-MLX-Serve-8bit). Unsloth lists 12-16 GB+ unified memory as a starting point. We found no measured Mac speed.

Does Ollama run Qwen-Image 2.1?

No. Ollama added experimental image generation on 20 January 2026. That post launched it on macOS, with Windows and Linux “coming soon.” On 23 September 2026, its image model namespace lists only x/z-image-turbo and x/flux2-klein. The URLs ollama.com/library/qwen-image and ollama.com/x/qwen-image-2.1 both return 404.

The name confusion is understandable. The files are called GGUF, and Ollama runs GGUF. But these GGUFs follow the stable-diffusion.cpp format. They carry no general.architecture metadata, which is why even the stock ComfyUI-GGUF node rejects them. Use ComfyUI or stable-diffusion.cpp instead.

Offloading and what it costs

  • Unsloth states FP8 runs on 6GB of VRAM with offloading, and that inference is “<2x slower.” That is a vendor claim, not a measurement we could check.
  • ComfyUI’s INT8 route on 16GB did not need an offload flag in the RTX 4060 Ti report.
  • stable-diffusion.cpp --offload-to-cpu made a 6GB card work at about 4 min 50 s per image. That is roughly 15x the 16GB card’s 20 s, but the two runs used different resolutions and runtimes.

Which GPU to Buy for It

Prices come from our tracked price file, as of August/September 2026.

  • Best fit: 16GB. The RTX 5060 Ti 16GB ($589-805 as of August 2026) matches the measured 16GB case: INT8 image model plus INT8 encoder, no offload, about 20 s per image on the older 4060 Ti.
  • Budget: 12GB. The RTX 3060 12GB ($329-460 new as of August 2026) holds the 9.91 GB Q4 GGUF stack. Expect slower images than 16GB. Nobody has published a 3060 timing yet.
  • Editing with reference images: 24GB. A used RTX 3090 24GB (about $1,453 average used as of September 2026) keeps 8-bit weights resident at 2048x2048 (22.6 GB measured peak). Multi-reference edits at 8-bit can exceed 24GB, so use 4-bit or offload for those.

Intel Arc and AMD cards: we found no measured Qwen-Image 2.1 report on either as of 23 September 2026. The community ncnn/Vulkan port claims NVIDIA, AMD, Intel and Apple support, but it is early. Buy NVIDIA for this model today. For other tiers, see the tested hardware list.

Honest Caveats

  1. We did not run the model. Every speed and peak number is a community report with a named source. Each is one GPU and one setup.
  2. The decoder table is arithmetic. It adds file sizes and an activation range taken from one RTX 5090 report. Real peaks vary with resolution, runtime and offload settings.
  3. Unsloth’s own table labels its figures “estimates, not tested minimums.” Its 24GB row suggests 512x512 for INT8/FP8, which is more cautious than the RTX 5090 measurement.
  4. The model is nine days old. ComfyUI support shipped two days ago. Expect new quants, fixes and speedups.
  5. Variants exist. Community fine-tunes and prompt-enhancer models (Qwen-Image-2.1-PE, based on Qwen3.5 9B, 9.47 GB at INT8) are separate downloads. The prompt enhancer is optional and adds its own memory cost.
  6. The license is non-commercial. Read the Qwen Research License before you use outputs in paid work.

Sources

See Also

Need OpenClaw fixed live?

Remote rescue sessions for gateway, auth, tunnel, VPS, and model access problems.

See Rescue Session

Read next

Can I Run LTX-2.5 Locally? VRAM by GPU (2026)
LTX-2.5 is a 22B video model plus a 12B Gemma 4 text encoder. The official INT8 files need a 24GB card. 16GB works with NVFP4 or GGUF and weight streaming. The '16GB minimum' leans on cloud text encoding. Ollama does not run it.
Can I Run MiniMax H3 Locally? Yes, 16GB Works (2026)
MiniMax H3 is a 33B audio-video model plus a Qwen3-VL-32B text encoder. A 16GB card runs ComfyUI's pruned INT8 files with offload. 24GB holds the video model. System RAM and the license territory matter more than most guides say.
Best Qwen Model for 16GB VRAM (2026): Qwen3.8 27B Wins
Best Qwen model for 16GB VRAM: Qwen3.8-27B at 3.5 bits (10.96 GiB). The Q4_K_M 27B files do not fit. Byte-exact GGUF sizes, KV math, and the best Qwen for 8, 12, 24 and 32GB.
Bonsai 2 27B on RTX 3060 12GB: Fits, Needs a Fork (2026)
Ternary Bonsai 2 27B is a 1.72-bit Qwen3.8-27B that fits a 12GB RTX 3060, and even an 8GB card. It needs the PrismML llama.cpp fork, not Ollama or LM Studio. File sizes, KV cache math for 12GB, vendor-reported quality, and how it compares to Qwen 3.8 27B on a 3090.