Can I Run Qwen-Image 2.1 Locally? Yes, on 12GB
Every Qwen-Image 2.1 GGUF page lists the image model alone: 4.20 GB at Q4_K_M. That number leaves out the bigger half. The model needs a Qwen3-VL 8B text encoder, and at BF16 that encoder is 17.53 GB, larger than the 14.23 GB image model. We read every file size from the Hugging Face API on 23 September 2026, checked ComfyUI, stable-diffusion.cpp and Ollama, and collected the measured VRAM reports with their sources.
Bottom Line
- Yes, you can run it locally. A 12GB card runs the Q4 GGUF stack. A 16GB card runs the INT8 stack without offload. A 6GB card runs it slowly with CPU offload.
- The GGUF size is not the VRAM answer. The Q4_K_M GGUF is 4.20 GB, but that file is the image model only. Add the text encoder and VAE and the Q4 stack is 9.91 GB.
- The text encoder is the biggest part. Qwen3-VL 8B is 17.53 GB at BF16. The 7B image model is 14.23 GB at BF16. Most guides skip this.
- Measured on 16GB: INT8 image model plus INT8 encoder peaked at about 15 GB on an RTX 4060 Ti 16GB, at about 20 s per image (community report).
- Ollama does not run it. Use ComfyUI (native since v0.37.0), diffusers, or stable-diffusion.cpp.
- License: Qwen Research License, non-commercial only.
Minimum VRAM by format, derived from the file sizes below:
| Format (image model + text encoder) | Files total | Minimum card |
|---|---|---|
| Q4_K_M GGUF + Q4_K_M encoder + VAE | 9.91 GB | 12GB (6-8GB with offload) |
| Q8_0 GGUF + Q4_K_M encoder + VAE | 13.35 GB | 16GB |
| INT8 + INT8 encoder + VAE (Comfy-Org) | 17.29 GB | 16GB (measured, ~15 GB peak) |
| 8-bit, all weights resident (diffusers) | 16.6 GB resident | 24GB (measured 21.3 GB peak) |
| BF16 + BF16 encoder + VAE | 32.44 GB | 32GB+ with offload, or 48GB |
What the Model Is
All figures come from the Qwen/Qwen-Image-2.1 model card, its config.json files and the Hugging Face API, read on 23 September 2026.
| Spec | Value |
|---|---|
| Image model | 7B parameters, 32 single-stream DiT layers, 32 heads x 128 dims |
| Text encoder | Qwen3-VL 8B (Qwen3VLForConditionalGeneration, 36 layers, hidden size 4096) |
| VAE | AutoencoderKLQwenImage21, its own VAE; earlier Qwen-Image VAEs do not work |
| Tasks | Text-to-image, image editing with up to 10 reference images, transparent RGBA output |
| Native resolution | 2048x2048 (1:1), up to 2752x1536 (16:9) |
| Default steps | 40 in the diffusers example |
| License | Qwen Research License Agreement, non-commercial |
| Created on Hugging Face | 14 September 2026 |
The model card says “7B parameters in its visual generation component.” The word “component” matters. The text encoder is a separate 8B vision-language model. It reads your prompt and any reference images before the image model starts.
Every File Size That Matters
Sizes come from the Hugging Face API with ?blobs=true, read on 23 September 2026. GB means 10^9 bytes.
Image model (the “diffusion model”)
| File | Source | Size (GB) |
|---|---|---|
| BF16 safetensors | Comfy-Org | 14.23 |
| INT8 ConvRot | Comfy-Org | 7.26 |
| FP8 | unsloth/Qwen-Image-2.1-FP8 | 7.12 |
| INT8 | unsloth/Qwen-Image-2.1-FP8 | 7.26 |
| F16 GGUF | unsloth/Qwen-Image-2.1-GGUF | 14.23 |
| Q8_0 GGUF | unsloth | 7.64 |
| Q6_K GGUF | unsloth | 6.27 |
| Q5_K_M GGUF | unsloth | 5.39 |
| Q4_K_M GGUF | unsloth | 4.20 |
| Q3_K_M GGUF | unsloth | 3.17 |
| Q2_K GGUF | unsloth | 2.47 |
| Q8_0 / Q6_K / Q4_K / Q2_K GGUF | leejet (stable-diffusion.cpp author) | 7.69 / 6.00 / 4.20 / 2.56 |
Text encoder (Qwen3-VL 8B)
| File | Source | Size (GB) |
|---|---|---|
| BF16 safetensors | Comfy-Org | 17.53 |
| INT8 ConvRot | Comfy-Org | 9.35 |
| W4A8 | Comfy-Org | 6.31 |
| FP8 | unsloth/Qwen-Image-2.1-FP8 | 9.39 |
| Q8_0 GGUF | unsloth/Qwen3-VL-8B-Instruct-GGUF | 8.71 |
| Q6_K GGUF | unsloth | 6.73 |
| UD-Q4_K_XL GGUF | unsloth (the encoder Unsloth recommends) | 5.15 |
| Q4_K_M GGUF | unsloth | 5.03 |
| mmproj F16 (vision part, needed for editing with a GGUF encoder) | unsloth | 1.16 |
VAE
The ComfyUI VAE file is 0.68 GB (qwen_image_2.1_vae_bf16.safetensors). One ComfyUI user reported an error with the VAE linked from the Unsloth card and fixed it with the official file (discussion).
The fact the GGUF pages leave out
At every precision, the text encoder is as big as the image model or bigger. At BF16 it is 17.53 GB against 14.23 GB. At 8 bits it is 9.35 GB against 7.26 GB. So “Q4_K_M = 4.20 GB” describes less than half of what you load. The Unsloth card does say the GGUF “is the denoiser only,” but the file list still leads with 4.20 GB.
There is a second cost. The model reuses a “prefix KV cache” for the prompt and reference images. The stable-diffusion.cpp docs state that a 4,096-token prefix takes about 4 GiB per condition in FP32. Positive and negative prompts keep separate caches. Reference images add prefix tokens, so editing costs more memory than plain text-to-image.
Decoder Table: Your GPU, Your Files
This table is our estimate. We added the file sizes above and checked the result against the measured reports in the next section. Activations at 1024x1024 add about 2-5 GB on top of resident weights, based on the RTX 5090 diffusers numbers (16.6 GB weights, 21.3 GB peak).
| VRAM | Image model | Text encoder | Expected experience |
|---|---|---|---|
| 6-8GB | Q4_0 or Q4_K_M GGUF (4.20 GB) | Q4_K_M GGUF (5.03 GB), offloaded to CPU | Works with --offload-to-cpu. Measured on a 6GB RTX 3050: about 4 min 50 s per 1184x1184 image, 25 steps. The tester had 16 GB of system RAM. |
| 12GB | Q4_K_M (4.20 GB) or Q6_K (6.27 GB) GGUF | Q4_K_M (5.03 GB) or UD-Q4_K_XL (5.15 GB) | Q4 stack is 9.91 GB, so it fits at 1024x1024, batch 1. Unsloth lists this tier. No published speed on a 12GB card yet. |
| 16GB | INT8 ConvRot (7.26 GB) or Q8_0 GGUF (7.64 GB) | INT8 ConvRot (9.35 GB) or W4A8 (6.31 GB) | Measured on an RTX 4060 Ti 16GB with INT8 files: ~15 GB peak, ~20 s per image, ~1 min per multi-image edit. |
| 24GB | INT8/FP8 (7.1-7.3 GB) or BF16 (14.23 GB) | INT8/FP8 (9.4 GB) | 8-bit with everything resident peaks at 21.3 GB (1024x1024) and 22.6 GB (2048x2048). A two-reference edit peaked at 25.6 GB, over 24GB. |
| 32GB | BF16 (14.23 GB) | INT8 (9.35 GB) or BF16 with offload | Holds the 8-bit two-reference edit (25.6 GB). The full BF16 stack is 32.44 GB of files, so BF16 everywhere still needs offload. |
Why the 16GB row works with 17.29 GB of files. The measured peak was about 15 GB with no offload flag. Our reading: ComfyUI runs the text encoder, then loads the image model, so both never sit in VRAM at full size together. The reporter did not state this; it is our inference from the numbers.
GGUF or INT8 on an NVIDIA card? An Unsloth team member wrote that GGUF “is naturally much slower and only meant for unified memory/CPU devices,” and that ComfyUI defaults to INT4 or INT8 (discussion). On a 16GB+ NVIDIA card, start with the Comfy-Org INT8 files. Use GGUF when VRAM forces it.
Measured Reports, With Sources
We found three measured reports as of 23 September 2026. We did not run these tests ourselves.
| GPU | Setup | Result | Source |
|---|---|---|---|
| RTX 5090 32GB | diffusers, fp8 (torchao) image model + encoder, all on GPU, 40 steps | Weights 16.6 GB; peak 21.3 GB at 1024x1024, 22.6 GB at 2048x2048; 19.3 s at 1024x1024, 115.7 s at 2048x2048 | HF discussion #32, 22 Sep 2026 |
| RTX 5090 32GB | diffusers, NF4 (bitsandbytes), all on GPU, 40 steps | Weights 10.6 GB; peak 15.2 GB at 1024x1024, 15.4 GB at 2048x2048; 19.2 s at 1024x1024 | same |
| RTX 4060 Ti 16GB | ComfyUI, INT8 ConvRot image model + INT8 Qwen3-VL 8B + BF16 VAE, no offload | Peak ~15 GB, typical 13-14 GB; ~20 s per image; ~1 min per multi-image edit | HF discussion #35, 22 Sep 2026 |
| RTX 3050 6GB, 16 GB RAM | stable-diffusion.cpp, Q4_0 GGUF + Q4_K_M encoder, --offload-to-cpu --diffusion-fa, CFG 1, 25 steps, 1184x1184 | ~4 min 50 s per image | leejet discussion #3, 23 Sep 2026 |
Three details from these reports change how you should run it:
- 2048x2048 costs 6x the time of 1024x1024. The RTX 5090 report measured 19.3 s against 115.7 s. Four times the pixels gave six times the time. Start at 1024x1024.
- Reference images cost VRAM. Two reference images raised the 8-bit peak from 21.3 GB to 25.6 GB. The 4-bit peak rose from 15.2 GB to 19.1 GB.
- CFG 1 halves the work. The RTX 3050 tester found CFG above 1 runs every step twice. CFG 1 was about twice as fast. The ComfyUI template also uses CFG 1.
--diffusion-facut step time by 35% on that card.
Which Tools Run It
We checked each runtime on 23 September 2026.
| Runtime | Status |
|---|---|
| ComfyUI | Native. PR #16400 “Qwen-image 2.1 support” merged 19 Sep 2026, shipped in v0.37.0 (21 Sep 2026). Official text-to-image and edit templates exist. |
| ComfyUI + GGUF | Needs a custom node. The stock city96 ComfyUI-GGUF node fails with “Unknown model architecture.” leejet recommends leejet/ComfyUI-GGUF; a community add-on also fixes it. |
| diffusers | Yes, QwenImage21Pipeline, from diffusers main (git install). Needs transformers>=5.17. enable_model_cpu_offload() is the official low-memory option. |
| stable-diffusion.cpp | Yes. Official docs page, GGUF files from leejet, --offload-to-cpu for small cards. |
| Unsloth Desktop | Yes, GGUF and FP8, text-to-image and editing. |
| Ollama | No. See below. |
| llama.cpp, LM Studio | Not listed by any model card we read. These are text-model runners. Do not assume support. |
| Apple Silicon | MLX conversions exist (for example ddalcu/Qwen-Image-2.1-MLX-Serve-8bit). Unsloth lists 12-16 GB+ unified memory as a starting point. We found no measured Mac speed. |
Does Ollama run Qwen-Image 2.1?
No. Ollama added experimental image generation on 20 January 2026. That post launched it on macOS, with Windows and Linux “coming soon.” On 23 September 2026, its image model namespace lists only x/z-image-turbo and x/flux2-klein. The URLs ollama.com/library/qwen-image and ollama.com/x/qwen-image-2.1 both return 404.
The name confusion is understandable. The files are called GGUF, and Ollama runs GGUF. But these GGUFs follow the stable-diffusion.cpp format. They carry no general.architecture metadata, which is why even the stock ComfyUI-GGUF node rejects them. Use ComfyUI or stable-diffusion.cpp instead.
Offloading and what it costs
- Unsloth states FP8 runs on 6GB of VRAM with offloading, and that inference is “<2x slower.” That is a vendor claim, not a measurement we could check.
- ComfyUI’s INT8 route on 16GB did not need an offload flag in the RTX 4060 Ti report.
- stable-diffusion.cpp
--offload-to-cpumade a 6GB card work at about 4 min 50 s per image. That is roughly 15x the 16GB card’s 20 s, but the two runs used different resolutions and runtimes.
Which GPU to Buy for It
Prices come from our tracked price file, as of August/September 2026.
- Best fit: 16GB. The RTX 5060 Ti 16GB ($589-805 as of August 2026) matches the measured 16GB case: INT8 image model plus INT8 encoder, no offload, about 20 s per image on the older 4060 Ti.
- Budget: 12GB. The RTX 3060 12GB ($329-460 new as of August 2026) holds the 9.91 GB Q4 GGUF stack. Expect slower images than 16GB. Nobody has published a 3060 timing yet.
- Editing with reference images: 24GB. A used RTX 3090 24GB (about $1,453 average used as of September 2026) keeps 8-bit weights resident at 2048x2048 (22.6 GB measured peak). Multi-reference edits at 8-bit can exceed 24GB, so use 4-bit or offload for those.
Intel Arc and AMD cards: we found no measured Qwen-Image 2.1 report on either as of 23 September 2026. The community ncnn/Vulkan port claims NVIDIA, AMD, Intel and Apple support, but it is early. Buy NVIDIA for this model today. For other tiers, see the tested hardware list.
Honest Caveats
- We did not run the model. Every speed and peak number is a community report with a named source. Each is one GPU and one setup.
- The decoder table is arithmetic. It adds file sizes and an activation range taken from one RTX 5090 report. Real peaks vary with resolution, runtime and offload settings.
- Unsloth’s own table labels its figures “estimates, not tested minimums.” Its 24GB row suggests 512x512 for INT8/FP8, which is more cautious than the RTX 5090 measurement.
- The model is nine days old. ComfyUI support shipped two days ago. Expect new quants, fixes and speedups.
- Variants exist. Community fine-tunes and prompt-enhancer models (Qwen-Image-2.1-PE, based on Qwen3.5 9B, 9.47 GB at INT8) are separate downloads. The prompt enhancer is optional and adds its own memory cost.
- The license is non-commercial. Read the Qwen Research License before you use outputs in paid work.
Sources
- Qwen/Qwen-Image-2.1 model card: 7B image model, 32 DiT layers, resolutions, steps, diffusers code, license
- Hugging Face API: Qwen/Qwen-Image-2.1: creation date, BF16 file sizes;
text_encoder/config.jsonandtransformer/config.jsonfor the architecture - Qwen Research License: non-commercial grant
- Comfy-Org/Qwen-Image-2.1: BF16, INT8 and W4A8 files, VAE, workflow templates
- unsloth/Qwen-Image-2.1-GGUF: GGUF sizes, “denoiser only” note, sd-cli command
- unsloth/Qwen-Image-2.1-FP8: FP8 and INT8 sizes
- unsloth/Qwen3-VL-8B-Instruct-GGUF: text encoder GGUF sizes
- leejet/Qwen-Image-2.1-GGUF: stable-diffusion.cpp GGUFs, ComfyUI-GGUF node advice
- Unsloth Qwen-Image-2.1 guide: memory tiers, 6GB offload claim, INT8 vs FP8 LPIPS
- stable-diffusion.cpp Qwen-Image 2.1 docs: prefix cache memory, VAE warning, commands
- ComfyUI PR #16400: native support, merged 19 Sep 2026
- Ollama image generation blog: experimental, supported models
- HF discussion #32: RTX 5090 fp8 and NF4 measurements
- HF discussion #35: RTX 4060 Ti 16GB INT8 measurements
- leejet discussion #3: RTX 3050 6GB stable-diffusion.cpp measurements
See Also
- Best local LLM for an RTX 3060 12GB: what else the 12GB budget card runs
- Best local LLM for an RTX 5060 Ti 16GB: the 16GB card this model fits best
- Best local LLM for an RTX 3090: the 24GB step for heavy editing work
- Is 16GB of VRAM still enough in 2026?: the tier-level verdict
- GGUF quant names explained: what Q4_K_M, Q6_K and Q8_0 mean
- Open text diffusion models you can run locally: the other kind of “diffusion” model
- Can I run LTX-2.5 locally?: the video-model version, where the Gemma 4 12B encoder is 37-43% of the files
- Can I run MiniMax H3 locally?: the video model that uses a Qwen3-VL-32B encoder, bigger than its pruned transformer
Need OpenClaw fixed live?
Remote rescue sessions for gateway, auth, tunnel, VPS, and model access problems.
See Rescue Session