← All guides

Best Qwen Model for 16GB VRAM (2026): Qwen3.8 27B Wins

On a 16GB card the best Qwen is Qwen3.8-27B, the newest dense model, at about 3.5 bits per weight. The 4-bit files most guides name do not fit: Qwen3.6-27B Q4_K_M is 15.66 GiB before any context. Here is the exact file for each VRAM tier, with byte counts from Hugging Face and the context each one leaves you.

Bottom Line

  • Best Qwen for 16GB VRAM: Qwen3.8-27B at about 3.5 bits. Use the ISTA-DASLab GSQ-RCO-IQ3_S file: 10.96 GiB, or 11.29 GiB with the MTP head.
  • It leaves room for 64K of context. With a Q8 KV cache, 64K costs 2.0 GiB. The total stays near 13.3 GiB plus buffers.
  • It keeps the quality. On its authors’ own run, IQ3_S scores 85.71 on LiveCodeBench v6, the same as BF16.
  • The popular 4-bit 27B files do not fit. Qwen3.6-27B Q4_K_M is 15.66 GiB. Qwen3.8-27B UD-Q4_K_XL is 16.35 GiB. Both fill the card before one token of context.
  • The MoE alternative is Qwen3.6-35B-A3B. It is faster and its cache is 3.2x smaller. It scores 9.9 points lower on LiveCodeBench v6 than Qwen3.8-27B.

Every file size on this page is a byte count from the Hugging Face API, read 2026-09-26. Every speed figure is community-reported, with the source named. We did not run them.

Ready to buy? See the tested hardware list with current prices.

The Current Qwen Lineup, and What It Means for a GPU Owner

We listed the Qwen organization on Hugging Face on 2026-09-26. Four facts decide the answer.

  1. Qwen3.8 has one model that fits a consumer card: the 27B. Qwen3.8-27B was created 2026-08-05. It has 27,781,427,952 parameters, an Apache 2.0 license and a 262,144-token native context.
  2. Qwen3.8-Flash-Next is not a small model. Its card lists 125B parameters with 6B active, plus a 51B n-gram embedding and a 4B MTP head. That is about 180B in total. It does not fit 16GB.
  3. Below 27B, the newest Qwen is still Qwen3.5. Qwen3.5-9B and Qwen3.5-4B date from February 2026. Qwen3.6 and Qwen3.8 have no 9B or 4B model.
  4. The only current Qwen MoE for a single GPU is the 35B-A3B. Qwen3.6-35B-A3B has 35B parameters with 3B active per token.

So a 16GB owner has three real choices: a dense 27B at low bits, the 35B-A3B MoE, or a 9B at high bits.

Best Qwen Model by VRAM Tier

VRAMBest QwenFileSizeContext that fitsSpeed (community-reported)
8GBQwen3.5-9BQ4_K_M (Unsloth)5.29 GiB32K (f16 cache 1.0 GiB)Not measured on an 8GB card
12GBQwen3.8-27BGSQ-RCO-IQ2_S (ISTA-DASLab)8.62 GiB32K (Q8 cache 1.0 GiB)No public figure for this file
16GBQwen3.8-27BGSQ-RCO-IQ3_S-mtp11.29 GiB64K (Q8 cache 2.0 GiB)29.18 tok/s, 63.89 with MTP, at 4K (similar-size UD-IQ3_XXS, RTX 5060 Ti)
24GBQwen3.8-27BUD-Q4_K_XL (Unsloth)16.35 GiB64K (f16 cache 4.0 GiB)40.31 tok/s at 4K (Q4_K_S, RTX 3090, Hardware Corner)
32GBQwen3.8-27BUD-Q6_K_XL (Unsloth)23.56 GiB128K (Q8 cache 4.0 GiB)about 100 tok/s with MTP at 120K (RTX 5090)

“Context that fits” is file size plus KV cache. It excludes compute buffers and the linear-attention state. Keep about 1 GiB free for those, and more if a monitor runs on the same card.

Speed sources: the 16GB row is from njannasch.dev (2026-08-15). The 24GB row is from Hardware Corner. The 32GB row is user PedroSantos76 in Unsloth discussion #14.

Why Qwen3.8-27B Wins at 16GB

It is the strongest Qwen you can hold on the card

Qwen’s own model card scores Qwen3.8-27B at 90.3 on LiveCodeBench v6 and 73.0 on Terminal Bench 2.1. The same card scores Qwen3.6-27B at 83.9 and 63.4. The Qwen3.6-35B-A3B card scores that MoE at 80.4 on LiveCodeBench v6.

Qwen3.6-27B appears at 83.9 on both the 3.6 card and the 3.8 card. So the two cards use a matching harness, and the comparison holds.

The 3.5-bit file keeps the benchmark score

ISTA-DASLab published its own results table for its GSQ-RCO files. Their harness gives BF16 a lower score than Qwen’s card does, so compare rows within their table only.

FileSizeLiveCodeBench v6GPQA-Diamond
BF16 (their run)53.8 GB (card)85.7189.90
GSQ-RCO IQ3_S10.96 GiB85.7189.39
GSQ-RCO IQ2_S8.62 GiB82.2986.36
Unsloth UD-IQ3_S11.21 GiB84.0089.90

At 3.5 bits the file ties BF16 on coding. It loses 0.51 points on GPQA-Diamond. We did not run these benchmarks. The full analysis is in our GSQ-RCO quants guide.

The KV cache is small for a 27B

Qwen3.8-27B uses hybrid attention. Its config.json lists 64 layers with full_attention_interval: 4. Only 16 layers keep a cache that grows with context.

2 (K and V) x 16 layers x 4 KV heads x 256 head_dim x 2 bytes (f16)
  = 65,536 bytes = 64 KiB per token
Q8 cache: about 32 KiB per token
ContextQ8 KV cache+ IQ3_S-mtp (11.29 GiB)Fits 16GB?
32K1.0 GiB12.29 GiBYes
64K2.0 GiB13.29 GiBYes
128K4.0 GiB15.29 GiBNo, no room for buffers

This is our arithmetic from the published config. Treat each row as a lower bound.

A real run shows how much the buffers add. User Bellatorius01 in Unsloth discussion #26 ran the 13.27 GiB UD-IQ4_XS file at 32K with a Q8 cache. The reported use was about 15.4 GB. File plus cache is 14.27 GiB, so buffers took roughly 1 GiB more. The IQ3_S-mtp file is 1.98 GiB smaller than that file.

Measured speed on a 16GB card

Nobody has published a speed figure for the GSQ-RCO IQ3_S file itself. Two community tests on an RTX 5060 Ti 16GB bracket it.

SourceFileSizeContextNormalWith MTP
njannasch.devUD-IQ3_XXS10.18 GiB4K29.18 tok/s63.89 tok/s
njannasch.devUD-IQ3_XXS10.18 GiB32K23.62 tok/s45.48 tok/s
njannasch.devUD-IQ3_XXS10.18 GiB64K19.77 tok/s39.06 tok/s
Bellatorius01, Unsloth #26UD-IQ4_XS13.27 GiBabout 2.5K26.0 tok/s54.1 tok/s

Our estimate for IQ3_S is 26 to 29 tok/s at short context without MTP. We derived it from the two files above, which sit either side of it in size. Reply speed on this card tracks file size, because the RTX 5060 Ti has 448 GB/s of bandwidth.

The njannasch test used a q4_0 KV cache and --spec-draft-n-max 3. It reported 15,732 MiB in use with a 96K prompt. That is the full card.

How to run it

hf download ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF \
  Qwen3.8-27B-GSQ-RCO-IQ3_S-mtp.gguf --local-dir .

llama-server -m Qwen3.8-27B-GSQ-RCO-IQ3_S-mtp.gguf \
  -ngl 99 -fa on -c 65536 --parallel 1 \
  --cache-type-k q8_0 --cache-type-v q8_0 \
  --spec-type draft-mtp --spec-draft-n-max 3

Skip the vision projector on 16GB. The mmproj file adds 0.87 GiB. For what each flag does, see llama.cpp flags explained.

Model and fileBytesGiBWhy it fails
Qwen3.5-27B Q4_K_M16,740,812,70415.59No room left for any cache
Qwen3.6-27B Q4_K_M16,817,244,38415.66No room left for any cache
Qwen3.8-27B UD-Q4_K_XL17,559,178,14416.35Larger than the card
Qwen3.6-35B-A3B UD-Q4_K_M22,134,528,99220.61Needs expert offload to system RAM
Qwen3.8-Flash-Nextabout 180B parametersn/aNeeds a multi-GPU server or a large Mac

The Q4_K_M 27B is the file most guides tell 16GB owners to download. A 16 GiB card gives you about 15.0 to 15.5 GiB after the display and the CUDA context. A 15.66 GiB file does not load with a usable context. Drop to 3.5 bits and the same model fits with 64K to spare.

The Real 16GB Decision: Dense 27B or the 35B-A3B MoE

Qwen3.6-35B-A3B activates 3B parameters per token, so it generates fast. Its config.json lists 40 layers, with full attention on 10 of them and 2 KV heads.

2 x 10 layers x 2 KV heads x 256 head_dim x 2 bytes = 20 KiB per token (f16)
128K context: 2.5 GiB   (Qwen3.8-27B: 8.0 GiB)

That gives you two ways to run the MoE on 16GB.

Option A: all on the GPU at 3 bits. The Unsloth UD-IQ3_XXS file is 12.30 GiB. Add 1.25 GiB of f16 cache for 64K and you reach 13.55 GiB. njannasch measured the earlier Qwen3.5-35B-A3B at the same quant on an RTX 5060 Ti: 47 to 51 tok/s at 160K context (source). Only 348 MiB of VRAM stayed free. Qwen3.5-35B-A3B and Qwen3.6-35B-A3B have identical layer configs. But the 3.6 file is 0.12 GiB larger than the 3.5 file (12.18 GiB), so plan on less than 160K. The author went back to a 9B for daily use because 348 MiB of headroom was too fragile.

Option B: full 4-bit with expert offload. Load the 20.61 GiB UD-Q4_K_M file. The --n-cpu-moe N flag keeps the expert weights of N layers in system RAM. Attention and shared weights stay on the GPU. InsiderLLM measured 38.2 tok/s at 8K this way on a 12GB RTX 3060, with -ncmoe 24 and 32GB of DDR4-2133 (source, firsthand). A 16GB card can keep more experts on the GPU. We found no single-source measurement on a 16GB card, so we give no figure.

llama-server -m Qwen3.6-35B-A3B-UD-Q4_K_M.gguf \
  -ngl 99 -fa on -c 65536 --n-cpu-moe 20
# Raise --n-cpu-moe until the model loads. Lower it until VRAM is full.
Qwen3.8-27B IQ3_S (dense)Qwen3.6-35B-A3B (MoE)
LiveCodeBench v6 (Qwen card)90.3 at BF1680.4 at BF16
KV cache per token (f16)64 KiB20 KiB
All on GPU at 16GBYes, 10.96 GiBOnly at 3 bits, 12.30 GiB
Needs 32GB system RAMNoYes, for the 4-bit file
Speed26-29 tok/s estimate, about 2x with MTP38-51 tok/s measured on 12-16GB cards

Pick the dense 27B for coding and agent work where a wrong answer costs you time. Pick the MoE when you need replies fast, or more than 64K of context on the card.

Tier Notes

8GB. Qwen3.5-9B Q4_K_M is 5.29 GiB. Its config lists 32 layers with 8 on full attention and 4 KV heads, so the f16 cache is 32 KiB per token. 32K of context adds 1.0 GiB. Qwen’s card scores it 65.6 on LiveCodeBench v6. More detail is in best local LLMs for 8GB.

12GB. Qwen3.8-27B GSQ-RCO IQ2_S is 8.62 GiB. With a Q8 cache, 32K adds 1.0 GiB. On ISTA-DASLab’s run it scores 82.29 on LiveCodeBench v6, against 85.71 for BF16. The fallback is Qwen3.5-9B at Q8_0, which is 8.87 GiB. The MoE with expert offload is the fast option: that is the 38.2 tok/s RTX 3060 result above.

To run the 16GB winner on an 8GB or 12GB system, the RTX 5060 Ti 16GB is the card both 16GB speed tests used. See our full 5060 Ti guide.

24GB. Qwen3.8-27B UD-Q4_K_XL is 16.35 GiB. 64K of f16 cache brings it to 20.35 GiB. This is the first tier where the 27B runs at 4 bits with no compromise. Measured speed and the MTP gain are in Qwen 3.8 27B on RTX 3090.

For a 16GB owner, 24GB is the only upgrade that changes the model. The used RTX 3090 24GB is the card in the Hardware Corner test. Used prices averaged about $1,463 as of September 2026.

32GB. Qwen3.8-27B UD-Q6_K_XL is 23.56 GiB. A Q8 cache at 128K adds 4.0 GiB, for 27.56 GiB. The 27B still wins here. Qwen3.8 has no model between 27B and the 180B Flash-Next.

FAQ

What is the best Qwen model for 16GB VRAM?

Qwen3.8-27B at about 3.5 bits per weight. The ISTA-DASLab GSQ-RCO IQ3_S file is 11,771,546,784 bytes (10.96 GiB). With a Q8 KV cache, 32K of context adds 1.0 GiB and 64K adds 2.0 GiB. On the authors' own run the file ties BF16 on LiveCodeBench v6 at 85.71. A community user measured a similar-size file (UD-IQ3_XXS) on an RTX 5060 Ti 16GB at 29.18 tok/s at 4K, and 63.89 tok/s with MTP.

Does Qwen 27B at Q4_K_M fit in 16GB VRAM?

No. Qwen3.6-27B Q4_K_M is 16,817,244,384 bytes, or 15.66 GiB. Qwen3.8-27B UD-Q4_K_XL is 16.35 GiB. Both fill a 16GB card before any KV cache or buffer. Use a 3 to 3.5 bit file of the same model, or buy 24GB.

What is the best Qwen model for 8GB VRAM?

Qwen3.5-9B at Q4_K_M, a 5.29 GiB file. Its KV cache costs 32 KiB per token at f16, so 32K of context adds 1.0 GiB. Qwen3.5 is still the newest Qwen text line below 27B. There is no small Qwen3.6 or Qwen3.8 model.

What is the best Qwen model for 24GB VRAM?

Qwen3.8-27B at UD-Q4_K_XL, a 16.35 GiB file. 64K of f16 context adds 4.0 GiB, for about 20.35 GiB. Hardware Corner measured 40.31 tok/s at 4K on an RTX 3090 with a Q4_K_S file of the same model.

Should I run Qwen3.6 35B-A3B or a dense 27B on 16GB?

Run the dense 27B for the best answers, and the 35B-A3B MoE for speed or very long context. Qwen's cards score Qwen3.8-27B at 90.3 on LiveCodeBench v6 and Qwen3.6-35B-A3B at 80.4. The MoE's cache is 20 KiB per token, so 128K costs 2.5 GiB against 8.0 GiB for the 27B. Its 4-bit file is 20.61 GiB and needs expert offload with --n-cpu-moe and 32GB of system RAM.

Sources

  • Hugging Face API, Qwen organization model list, sorted by creation date, read 2026-09-26
  • Qwen/Qwen3.8-27B, Qwen/Qwen3.6-27B, Qwen/Qwen3.6-35B-A3B, Qwen/Qwen3.5-35B-A3B, Qwen/Qwen3.5-9B, Qwen/Qwen3.5-4B: API metadata, config.json and README benchmark tables, read 2026-09-26
  • Qwen/Qwen3.8-Flash-Next README (“125B with 6B activated, plus 51B n-gram embedding and 4B MTP”) and API parameter count, read 2026-09-26
  • Hugging Face API ?blobs=true file sizes: unsloth/Qwen3.8-27B-GGUF, unsloth/Qwen3.6-27B-GGUF, unsloth/Qwen3.6-35B-A3B-GGUF, unsloth/Qwen3.5-27B-GGUF, unsloth/Qwen3.5-9B-GGUF, ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF, read 2026-09-26
  • njannasch.dev, “Qwen 3.8 27B on a 5060 Ti” (2026-08-15) and “Running Qwen 3.5 35B-A3B on a 5060 Ti” (2026-03-02), community-reported
  • Unsloth Hugging Face discussions #14 (RTX 5090) and #26 (RTX 5060 Ti), community-reported
  • Hardware Corner, “We Tested Qwen3.8 27B” (llama.cpp build 153d324bc, MTP off)
  • InsiderLLM, “Best Way to Run Qwen 3.6 35B MoE Locally” (updated 2026-09-22), firsthand RTX 3060 result
  • OpenClaw DC hardware price reference, used RTX 3090 figure re-read 2026-09-24

See Also

Need OpenClaw fixed live?

Remote rescue sessions for gateway, auth, tunnel, VPS, and model access problems.

See Rescue Session

Read next

Can I Run MiniMax H3 Locally? Yes, 16GB Works (2026)
MiniMax H3 is a 33B audio-video model plus a Qwen3-VL-32B text encoder. A 16GB card runs ComfyUI's pruned INT8 files with offload. 24GB holds the video model. System RAM and the license territory matter more than most guides say.
Can I Run LTX-2.5 Locally? VRAM by GPU (2026)
LTX-2.5 is a 22B video model plus a 12B Gemma 4 text encoder. The official INT8 files need a 24GB card. 16GB works with NVFP4 or GGUF and weight streaming. The '16GB minimum' leans on cloud text encoding. Ollama does not run it.
Can I Run Qwen-Image 2.1 Locally? Yes, on 12GB
Qwen-Image 2.1 is a 7B image model, but its Qwen3-VL 8B text encoder is bigger than the image model. A Q4 GGUF stack is 9.91 GB, so 12GB works. 16GB runs INT8 at about 20 s per image. Ollama does not run it.
Bonsai 2 27B on RTX 3060 12GB: Fits, Needs a Fork (2026)
Ternary Bonsai 2 27B is a 1.72-bit Qwen3.8-27B that fits a 12GB RTX 3060, and even an 8GB card. It needs the PrismML llama.cpp fork, not Ollama or LM Studio. File sizes, KV cache math for 12GB, vendor-reported quality, and how it compares to Qwen 3.8 27B on a 3090.