← All guides

Best Local LLM for R9700 (2026): Qwen3.6 35B-A3B Wins

The AMD Radeon AI PRO R9700 has 32GB of GDDR6 at 640 GB/s. That is the capacity of an RTX 5090 at about a third of its bandwidth. So the best model for this card is a mixture-of-experts model: Qwen3.6 35B-A3B at Q6_K fits with a 128K context and reads only 3B parameters per token. This page lists what fits, the community-measured speeds under ROCm and Vulkan, and where the card loses to CUDA.

Shopping? See the tested hardware list with every verified pick by tier.

Bottom Line

  • Best model for the R9700: Qwen3.6 35B-A3B at Q6_K. The file is 27.3 GiB. It fits with a 128K context. Only 3B parameters are active per token, so the card’s 640 GB/s goes far.
  • Measured speed for this model class: 77 tok/s on ROCm (Qwen3.6 35B-A3B, Q5_K_M) and 127-150 tok/s on Vulkan (Qwen3.5 35B-A3B, Q4_K_XL). Both are community-reported.
  • Best dense model: Qwen3.8 27B at Q6_K (20.5 GiB, 128K context). Expect about 30 tok/s. Turn on MTP under ROCm for 46 tok/s.
  • Best for agent tool calls: gpt-oss 20B. It is 12.9 GiB and holds its full 128K context with about 14 GiB to spare.
  • Pick the backend per job. Vulkan decodes MoE models up to about twice as fast as ROCm. ROCm prefills a dense 27B prompt about 5x faster than Vulkan.
  • Price: about $1,599 street average (range $1,400-1,900), as of September 2026. The $1,299 MSRP listing at Newegg was out of stock on 2026-09-26.

The R9700 is a used-RTX-3090-speed card with 32GB. On the llama.cpp scoreboards it decodes Llama 2 7B at 152.7 tok/s on Vulkan. A used RTX 3090 does 161.9 tok/s on CUDA. The R9700 has 8GB more memory and a warranty. It loses to CUDA on prompt processing and on software breadth.

R9700 VRAM and bandwidth

SpecificationRadeon AI PRO R9700
Memory32GB GDDR6
Bus width256-bit
Bandwidth640 GB/s
Compute units64 (RDNA4, gfx1201)
Infinity Cache64MB
Total board power300W
Minimum PSU (AMD)750W
InterfacePCIe 5.0 x16

Read from AMD’s R9700 specification page on 2026-09-26.

Bandwidth sets decode speed. A dense model reads all its weights for each token. So a 20 GiB dense file on a 640 GB/s card has a decode ceiling near 30 tok/s (640 ÷ ~21.5 GB). That is our arithmetic, not a measurement. A MoE model with 3B active parameters reads a small part of its file for each token. That is why MoE models win on this card.

Best Local LLMs for the R9700

1. Qwen3.6 35B-A3B (Q6_K): the winner

Qwen3.6 35B-A3B has 36B total parameters and 256 experts, with 8 active per token. Only 10 of its 40 layers use full attention. That makes its KV cache small: about 20 KiB per token at FP16, by our calculation from the model config. A 128K context costs about 2.5 GiB.

QuantFile sizeContext that fits (FP16 KV)Measured on R9700
UD-Q4_K_XL20.8 GiBFull 262K127-150 tok/s, Vulkan (Qwen3.5 35B-A3B, same size class)
UD-Q5_K_M24.6 GiBFull 262K77.25 tok/s, ROCm
Q6_K27.3 GiB128KNot measured
Q8_034.4 GiBDoes not fit one cardTwo cards only

The Vulkan figures come from two llama.cpp discussions on Qwen3.5 35B-A3B UD-Q4_K_XL: 127.4 tok/s (discussion #19890) and 147.8-149.5 tok/s after tuning (discussion #21043). The ROCm figure is a single-card run of Qwen3.6 35B-A3B UD-Q5_K_M on ROCm 7.2.2. All three are community-reported.

# Vulkan build: the fast path for MoE decode
cmake -B build -DGGML_VULKAN=ON && cmake --build build -j
./build/bin/llama-server -hf unsloth/Qwen3.6-35B-A3B-GGUF:Q6_K -ngl 99 -c 131072

The same model is the pick on our RTX 5090 page. The 5090 runs it faster. The R9700 runs the same file at about a third of the card price.

2. Qwen3.8 27B (Q6_K): best dense model

Qwen3.8 27B is a dense 27.8B model from August 2026. It has 64 layers, and 16 use full attention. Its KV cache costs about 64 KiB per token at FP16, so 128K costs 8 GiB.

  • Q6_K is 20.5 GiB. It fits with 128K context.
  • UD-Q4_K_XL is 16.4 GiB. It fits 128K at FP16 KV, or the full 262K with a q8_0 KV cache.
  • Q8_0 is 27.1 GiB. It leaves room for about 48K of context at FP16 KV (our estimate).

Measured, community-reported: 29.8 tok/s on ROCm without speculation and 46.2 tok/s with MTP (draft 2), on IQ4_NL at 131K context (akougkas.io, 2026-09-08). AMD’s own blog reports up to 51.8 tok/s with MTP=2 on Windows Vulkan. Another tuned run measured 64.95 tok/s on Q8_0 with MTP (llama.cpp discussion #21043).

On a 24GB RTX 3090 the same model does 41.55 tok/s at 8K. See Qwen 3.8 27B on RTX 3090. The 3090 is faster. The R9700 holds a higher quant with a longer context.

3. gpt-oss 20B: best for agent tool calls

The native MXFP4 file is 12.9 GiB. Its full 128K context costs 3.0 GiB. So the model and its whole context use under half the card. We found prefill figures for the R9700 (1,323-3,674 tok/s, by batch size, Vulkan) but no measured decode figure. We do not quote one.

4. Other 27-31B dense models

Measured on one R9700 with ROCm 7.2.2 (community-reported):

  • Qwen3.6 27B Q5_K_M (18.2 GiB): 24.85 tok/s decode, 611 tok/s prefill at 32K.
  • Gemma 4 31B Q5_K_M (20.2 GiB): 21.74 tok/s decode, 425 tok/s prefill at 32K.

Qwen3.8 27B replaces Qwen3.6 27B at the same size. Gemma 4 31B is slower on this card.

What fits in 32GB on the R9700

File sizes from the Hugging Face API, read 2026-09-26. Context budgets are our arithmetic: we allow 30 GiB for weights plus KV cache and keep about 2 GiB for compute buffers.

ModelQuantFile sizeFits one R9700?Context at FP16 KVMeasured decode (source)
Qwen3.6 35B-A3BQ6_K27.3 GiBYes128Kn/a
Qwen3.6 35B-A3BUD-Q5_K_M24.6 GiBYes262K77.25 tok/s ROCm
Qwen3.5 35B-A3BUD-Q4_K_XL~18.3 GiBYes262K127-150 tok/s Vulkan
Qwen3.8 27BQ6_K20.5 GiBYes128Kn/a
Qwen3.8 27BIQ4_NL15.2 GiBYes131K tested29.8 ROCm, 46.2 ROCm + MTP
Qwen3.8 27BQ8_027.1 GiBYes~48K (estimate)64.95 Vulkan + MTP
gpt-oss 20BMXFP412.9 GiBYesFull 128Kn/a
Gemma 4 31BQ5_K_M20.2 GiBYesNot calculated21.74 tok/s ROCm
gpt-oss 120BMXFP459.0 GiBNon/aTwo cards, tight
Qwen3.8 Flash-Next (180B)UD-IQ1_S67.6 GiBNon/aNot even on two cards

The Qwen3.5 file size is the 18.32 GiB that the benchmark author reports.

ROCm or Vulkan on the R9700

No single backend wins on this card. The numbers, all from llama.cpp:

TestROCm (HIP)VulkanSource
Llama 2 7B Q4_0, decode (FA on)97.98 tok/s152.70 tok/sllama.cpp scoreboards #15021 and #10879
Llama 2 7B Q4_0, prefill pp512 (FA on)4,773 tok/s5,908 tok/sSame
35B-A3B MoE, decode77.25 tok/s (Qwen3.6, Q5_K_M)127-150 tok/s (Qwen3.5, Q4_K_XL)truelies444 repo; #19890, #21043
Qwen3.8 27B IQ4_NL, decode29.8 tok/s32.9 tok/sakougkas.io
Qwen3.8 27B IQ4_NL, prefill1,287 tok/s239 tok/sakougkas.io
Qwen3.8 27B with MTP, decode46.2 tok/s28.0 tok/sakougkas.io

The MoE row compares two model versions at two quants. Treat it as a direction, not an exact ratio. The MTP row is one Linux setup; see rule 3.

The rule:

  1. Chat with a MoE model: use Vulkan.
  2. Dense model with long prompts (coding agents, RAG): use ROCm. Vulkan prefill on the dense 27B was 5x slower.
  3. MTP speculative decoding: test both backends. In one Linux test, MTP on Vulkan made Qwen3.8 27B slower (28.0 against 32.9 tok/s) and cut prefill to 74 tok/s. ROCm rose to 46.2. Other Vulkan runs gained from MTP: AMD measured 51.8 tok/s on Windows, and discussion #21043 measured 64.95 tok/s on Q8_0.
# ROCm build for the R9700 (Linux)
cmake -B build -DGGML_HIP=ON -DAMDGPU_TARGETS=gfx1201 && cmake --build build -j

ROCm support is official. AMD’s ROCm 7.14 system requirements list the R9700 as gfx1201. For Radeon cards, ROCm supports Ubuntu 24.04.4, Ubuntu 22.04.5, RHEL 10.1 and RHEL 9.7 only. Ollama lists the R9700 under ROCm support. On Windows, use Vulkan. AMD measured its own 51.8 tok/s figure that way.

Where the R9700 loses to CUDA cards

TestR9700RTX 3090 (CUDA)RTX 5090 (CUDA)Source
Llama 2 7B Q4_0 decode (FA on)152.7 (Vulkan)161.9300.4llama.cpp scoreboards
Llama 2 7B Q4_0 prefill pp512 (FA on)5,908 (Vulkan)5,56014,970Same
Qwen3.5 35B-A3B prefill at 5122,713 (Vulkan)n/a7,026Discussion #19890
Qwen3.5 35B-A3B prefill at 32K1,877 (Vulkan)n/a6,461Same
Qwen3.5 35B-A3B decode127.4 (Vulkan)n/a194.0Same

What you give up:

  • Prefill on long prompts. On the Qwen 35B-A3B test, the R9700’s prefill fell 31% from 512 to 32K tokens. The 5090 fell 8%. At 32K the 5090 is 3.4x faster. A coding agent that re-reads a large repo feels this gap.
  • One-click serving. vLLM lists gfx1200/gfx1201 as supported. Puget Systems ran stock vLLM with two-card tensor parallelism on bare metal. It failed in virtual machines. The fastest dual-R9700 vLLM result we found used a community-patched build.
  • Operating system choice. ROCm runs on four Linux releases. Windows users get Vulkan only.
  • Tooling breadth. Most local AI tools test on CUDA first. We did not test image or video generation for this page.

The R9700 still wins on one axis: 32GB. A used 3090 cannot hold Qwen3.6 35B-A3B at Q6_K. An RTX 5090 can, at a $4,299 floor (as of September 2026).

Dual R9700 (64GB): the next tier

Two R9700s give 64GB. Community tests show what the second card does and does not do.

Test (ROCm 7.2.2)One R9700Two R9700s
Qwen3.6 27B Q5_K_M, prefill at 32K611 tok/s1,216 tok/s
Qwen3.6 27B Q5_K_M, decode24.85 tok/s24.31 tok/s
Qwen3.6 35B-A3B UD-Q5_K_M, prefill at 32K1,637 tok/s3,038 tok/s
Qwen3.6 35B-A3B UD-Q5_K_M, decode77.25 tok/s71.88 tok/s

Source: truelies444 dual-R9700 benchmark repo, llama.cpp with ROCm 7.2.2.

A second card doubles long-prompt speed. It does not make single-user chat faster. For several users it helps. Puget measured stock vLLM on two cards at 156.2 tok/s total across 8 users, on a 27B model in FP8.

What 64GB adds:

  • Qwen3.6 35B-A3B at Q8_0 (34.4 GiB) with room for long context.
  • gpt-oss 120B (59.0 GiB). This fit is tight. We found no two-card measurement. One user ran it on three R9700s under vLLM at 21.48 tok/s.

Two cards at a 300W board power each need a larger supply than AMD’s 750W minimum for one card. A 1200W unit such as the MSI MAG A1200PLS 1200W leaves headroom. See what PSU for a local AI rig for the sizing math. At the $1,599 street average, two cards cost about $3,198 (as of September 2026). One RTX 5090 costs $4,299 at the floor and has half the memory.

R9700 vs the cards you are also looking at

Prices as of September 2026.

CardPriceVRAMBandwidthHolds Qwen3.6 35B-A3B Q6_K (27.3 GiB)?
Radeon AI PRO R9700~$1,599 street avg ($1,400-1,900)32GB640 GB/sYes, with 128K
Used RTX 3090~$1,463 avg24GB936 GB/sNo
RTX 5090$4,299 floor32GB1,792 GB/sYes
  • Buy the R9700 if you need 32GB on one card and can run Linux or Vulkan. The XFX Radeon AI PRO R9700 32GB is the catalog pick. Compare the price against the $1,400-1,900 street range before you pay.
  • Buy a used RTX 3090 if your models fit in 24GB and you want CUDA. It decodes about as fast. The EVGA RTX 3090 24GB is the used-market alternative.
  • Buy an RTX 5090 if long-prompt prefill is your bottleneck. It is 2.6-3.4x faster at prefill on the same MoE model. The GIGABYTE RTX 5090 32GB costs 2.7x the R9700’s street average.

Common mistakes on the R9700

  1. Using one backend for everything. Vulkan wins MoE decode. ROCm wins dense prefill.
  2. Assuming MTP always helps. In one Linux Vulkan test it lowered Qwen3.8 27B decode from 32.9 to 28.0 tok/s. Measure with and without it.
  3. Buying a second card for chat speed. Two cards doubled long-prompt prefill. Single-user decode stayed flat.
  4. Picking a dense 27B for speed. At 640 GB/s, a dense 27B decodes near 25-33 tok/s. A 35B-A3B MoE decodes at 77-150.
  5. Installing ROCm on an unlisted distro. ROCm supports the R9700 on Ubuntu 24.04.4, 22.04.5, RHEL 10.1 and RHEL 9.7.

FAQ

Is the R9700 good for local LLMs?

Yes, for models that fit in 32GB, and best for mixture-of-experts models. The R9700 has 32GB of GDDR6 at 640 GB/s. On the llama.cpp Vulkan scoreboard it decodes Llama 2 7B Q4_0 at 152.7 tok/s, close to a used RTX 3090 on CUDA (161.9 tok/s), with 8GB more memory. It is slower than an RTX 5090 on prompt processing: 2,713 against 7,026 tok/s on a community Qwen3.5 35B-A3B test. As of September 2026 its street price averages about $1,599.

What LLM can I run on an R9700?

Qwen3.6 35B-A3B at Q6_K (27.3 GiB) with a 128K context is the best pick. Community tests of this model class measure 77 tok/s on ROCm and 127-150 tok/s on Vulkan. Qwen3.8 27B fits at Q6_K (20.5 GiB) with 128K context and decodes at about 30 tok/s, or 46 tok/s with MTP on ROCm. gpt-oss 20B fits with its full 128K context. gpt-oss 120B (59 GiB) and Qwen3.8 Flash-Next (67.6 GiB at its smallest quant) do not fit on one card.

Does the R9700 work with ROCm?

Yes. AMD's ROCm 7.14 system requirements list the Radeon AI PRO R9700 as RDNA4, LLVM target gfx1201, supported. For Radeon cards, ROCm supports only Ubuntu 24.04.4, Ubuntu 22.04.5, RHEL 10.1 and RHEL 9.7. Ollama lists the R9700 under ROCm support, and vLLM lists gfx1201. On Windows, use the Vulkan backend of llama.cpp or LM Studio.

How much VRAM and memory bandwidth does the R9700 have?

32GB of GDDR6 on a 256-bit bus at 640 GB/s, per AMD's specification page. It has 64 compute units, 64MB of Infinity Cache, a 300W total board power and a PCIe 5.0 x16 interface. AMD recommends a 750W power supply for one card. The bandwidth is 36% of an RTX 5090's 1,792 GB/s and 68% of an RTX 3090's 936 GB/s.

Is a dual R9700 (64GB) setup worth it?

Only for long prompts, bigger quants, or several users. In a community test on two R9700s with ROCm, prefill at 32K roughly doubled (611 to 1,216 tok/s on Qwen3.6 27B), but single-user decode did not rise (24.85 to 24.31 tok/s). Puget Systems ran stock vLLM with tensor parallelism on two R9700s on bare metal and measured 156.2 tok/s total across 8 users. gpt-oss 120B (59 GiB) is a tight fit on 64GB.

Sources

Before you order parts, check the tested hardware list for current prices by tier.

See Also

Need OpenClaw fixed live?

Remote rescue sessions for gateway, auth, tunnel, VPS, and model access problems.

See Rescue Session

Read next

Best Local LLM for 64GB VRAM (2026): Laguna S 2.1 Wins
Best local LLM for 64GB VRAM: Laguna S 2.1 UD-IQ4_XS (53.6 GiB) split over two RTX 5090s, Qwen3.8 27B Q8 on one card. Why gpt-oss 120B does not fit.
Best Local LLM for 64GB RAM (2026): gpt-oss 120B Wins
Best local LLMs for 64GB RAM in 2026. Llama 4 Scout (10M context, ~58GB Q4), gpt-oss 120B at Q4, DeepSeek V4 Flash (284B MoE, Ollama cloud), Laguna XS 2.1 (agentic coding, 33B-A3B, ~36GB Q8). Also: Mistral Small 4, Qwen 3.6 35B Q8.
Best Local LLM for RTX PRO 5000 (2026): gpt-oss 120B Wins
What runs on the RTX PRO 5000 Blackwell 48GB and 72GB: gpt-oss 120B with full 128K context on the 72GB, Qwen3.6 35B-A3B at Q8 on the 48GB, measured tok/s, $/GB.
Is the RTX PRO 4500 Worth $4,700 for Local LLMs? (2026)
The RTX PRO 4500 Blackwell has the 32GB of an RTX 5090 at half the bandwidth and 200W. In stock it costs $4,600-5,285 as of September 2026, more than a $4,299 RTX 5090. A verdict per buyer, $ per GB/s, community llama.cpp speed, and the one build where it wins.