← All guides

Best Local LLM for RTX PRO 5000 (2026): gpt-oss 120B Wins

The RTX PRO 5000 Blackwell ships in two sizes: 48GB and 72GB of GDDR7, both at 1,344 GB/s. The 72GB card is the smallest single NVIDIA card that holds gpt-oss 120B with its full 128K context. The 48GB card misses it by 11 GiB, so its best model is Qwen3.6 35B-A3B at Q8_0. This page lists what fits on each variant, the measured speeds, and what two RTX 5090s or an RTX PRO 6000 give you for the same money.

Shopping? See the tested hardware list with every verified pick by tier.

Bottom Line

  • Best model on the 72GB card: gpt-oss 120B. The MXFP4 file is 59.03 GiB. Its full 128K context costs about 4.5 GiB. The whole job fits on one card with room left.
  • Measured speed: Hostkey reports 156.52 tok/s for gpt-oss 120B on the 72GB card. This is community-reported, and the runtime is not stated.
  • Best model on the 48GB card: Qwen3.6 35B-A3B at Q8_0. The file is 34.37 GiB. It fits with the full 262K context. Only 3B parameters are active per token.
  • The 48GB card cannot hold gpt-oss 120B. It reports 47.8 GiB usable. The weights alone are 11.2 GiB too large.
  • Both variants have the same speed. They share 1,344 GB/s, 14,080 CUDA cores and 300W. The 72GB adds capacity only.
  • Price, as of September 2026: 48GB at $7,499.99 and 72GB at $8,999.99 at Central Computer. The 72GB was listed as unavailable online.

The 72GB card is cheaper per gigabyte than the 48GB card. At the one US retailer that lists both, it costs $125/GB against $156/GB. It also undercuts the RTX PRO 6000 ($167/GB) and two RTX 5090s at their best deal price ($134/GB).

RTX PRO 5000 VRAM and bandwidth

SpecificationRTX PRO 5000 48GBRTX PRO 5000 72GB
Memory48GB GDDR7 ECC72GB GDDR7 ECC
Bus width384-bit384-bit
Bandwidth1,344 GB/s1,344 GB/s
CUDA cores14,08014,080
Max power300W300W
CoolerDual slot, activeDual slot, active
MIGUp to 2 instancesUp to 2 x 36GB
Usable in nvidia-smi48,935 MiB (47.8 GiB), measured~71.7 GiB (our estimate)

Specifications from NVIDIA’s RTX PRO 5000 page and PNY’s 72GB product page, read 2026-09-27. The 48GB nvidia-smi figure comes from a ComputingForGeeks test (2026-09-25). We found no nvidia-smi reading for the 72GB card. Our 71.7 GiB figure scales the 48GB reading by 1.5.

Decode speed follows bandwidth. For each token, a dense model reads all its weights. A MoE model reads only its active experts. So on a 1,344 GB/s card, MoE models are the fast choice.

Best Local LLMs for the RTX PRO 5000 72GB

1. gpt-oss 120B (MXFP4): the winner

gpt-oss 120B has 117B parameters with 5.1B active per token, per OpenAI’s model card. The native MXFP4 GGUF (ggml-org) is 63,387,346,208 bytes, which is 59.03 GiB.

Its KV cache is small. Half of its 36 layers use a 128-token sliding window. The other 18 layers cost 36 KiB per token at FP16. So the full 131,072-token window costs about 4.5 GiB. See gpt-oss 120B VRAM at full context for the arithmetic.

ItemSize
Weights, MXFP4 GGUF59.03 GiB
KV cache, FP16, full 131K4.5 GiB
Total63.5 GiB
Left on the 72GB card~8 GiB (estimate) for buffers
./build/bin/llama-server -hf ggml-org/gpt-oss-120b-GGUF -ngl 99 -c 131072 --jinja

Measured: 156.52 tok/s (Hostkey, 72GB card, 2026-09-08). Hostkey does not name its runtime or context length. Treat the number as one data point.

The line other reviews miss: the context is not what blocks the 48GB card. The full 128K cache is 4.5 GiB. The 48GB card is 11.2 GiB short on the weights alone. No context setting fixes that.

2. Two models resident at once

The 72GB card can hold Qwen3.6 35B-A3B at Q8_0 (34.37 GiB) and Qwen3.8 27B at Q8_0 (27.05 GiB) together. That is 61.4 GiB of weights. With 64K of context each, the caches add 1.25 GiB and 4.0 GiB. The total is about 66.7 GiB. This fit is tight. It is our arithmetic, not a measurement.

This is the agent setup: a fast MoE model for tool calls and a dense 27B for careful code.

3. Qwen3.8 27B at BF16: an unquantized 27B

The BF16 file is 50.90 GiB. With a 128K context (8 GiB), it uses about 58.9 GiB. You get the model with no quantization loss. The cost is speed: a dense 55 GB read per token.

Best Local LLMs for the RTX PRO 5000 48GB

1. Qwen3.6 35B-A3B (Q8_0): the winner

Qwen3.6 35B-A3B has 35B total parameters and 3B active, with 256 experts. Only 10 of its 40 layers use full attention. From the config (2 KV heads, head_dim 256), that is 20 KiB per token at FP16. Its native 262,144-token context costs 5.0 GiB.

At Q8_0 (34.37 GiB) plus the full 262K cache, the total is 39.4 GiB. That leaves about 8 GiB on the 48GB card. You run the model at 8-bit with its whole window.

2. Qwen3.8 27B (Q8_0): best dense model

Qwen3.8 27B has 64 layers, and 16 use full attention (4 KV heads, head_dim 256). That is 64 KiB per token, so 262K costs 16 GiB. The Q8_0 file is 27.05 GiB. Model plus the full window is 43.1 GiB. It fits on the 48GB card, with about 4.7 GiB left.

3. gpt-oss 20B: the lightweight agent model

gpt-oss 20B has 21B parameters and 3.6B active. The unsloth F16 file, with MXFP4 experts, is 12.85 GiB. Its full 131K context costs 3.0 GiB. It uses under a third of the card.

What about gpt-oss 120B on 48GB?

It does not fit on the GPU. You can keep some expert layers in system RAM with llama.cpp’s --n-cpu-moe option. Speed then depends on your CPU memory. We found no measurement on this card, so we give no number.

What fits on each variant

File sizes from the Hugging Face API, read 2026-09-27. Context budgets are our arithmetic from each model’s config.json at FP16 KV. We keep 2 GiB free for compute buffers.

ModelQuantFile48GB card72GB cardDecode (label)
gpt-oss 120BMXFP459.03 GiBNoYes, full 131K156.52 tok/s, measured (Hostkey)
Qwen3.6 35B-A3BQ8_034.37 GiBYes, full 262KYes, full 262K~140 tok/s, estimate
Qwen3.8 27BQ8_027.05 GiBYes, full 262KYes, full 262K19-37 tok/s, estimate
Qwen3.8 27BBF1650.90 GiBNoYes, 128K10-19 tok/s, estimate
gpt-oss 20BF16 (MXFP4 experts)12.85 GiBYes, full 131KYes, full 131K~185 tok/s, estimate
Qwen3.6 35B-A3B + Qwen3.8 27BQ8_0 + Q8_061.42 GiBNoYes, 64K each (tight)n/a
Qwen3.8 Flash-NextUD-IQ1_S (smallest)67.56 GiBNoLoads, almost no contextNot recommended
DeepSeek V4 Flash 0731UD-IQ1_S (smallest)76.87 GiBNoNon/a
GLM-5.3 FlashUD-IQ1_S (smallest)86.69 GiBNoNon/a

How we derived the estimates. We read the ceiling as 1,344 GB/s divided by the bytes read per token. Then we apply the efficiency this card reached in published tests:

  • MoE models, about 33% of the ceiling. Hostkey’s 156.52 tok/s on gpt-oss 120B is about a third of its ceiling. That ceiling is 1,344 ÷ ~2.85 GB of active weights per token, or ~470 tok/s.
  • Qwen3.6 35B-A3B Q8_0: 36.9 GB x 3/35 = 3.16 GB per token. 1,344 ÷ 3.16 = 425 tok/s ceiling. At 33%, about 140 tok/s.
  • gpt-oss 20B: 13.8 GB x 3.6/21 = 2.37 GB per token. Ceiling 568 tok/s. At 33%, about 185 tok/s.
  • Dense models, 41-79% of the ceiling. That is the range between two published tests (see below).
  • Qwen3.8 27B Q8_0: 29.05 GB per token. Ceiling 46.3 tok/s. At 41-79%, 19-37 tok/s.

These are estimates, not measurements. Your runtime, context length and settings move them.

Measured results on this chip

Model (file size)CardDecodeSource
gpt-oss 120B72GB156.52 tok/sHostkey, 2026-09-08
Qwen3-Coder-Next q4_K_M72GB192.13 tok/sHostkey
DeepSeek-R1 70B (42.52 GB)72GB25.08 tok/sHostkey
DeepSeek-R1 32B (19.85 GB)72GB50.77 tok/sHostkey
Qwen2.5 32B Q4_K_M (19.85 GB)48GB27.8 tok/sComputingForGeeks, Ollama, 2026-09-25
Qwen2.5 14B Q4_K_M48GB54.7 tok/sComputingForGeeks, Ollama
Llama 3.1 8B Q4_K_M (4.92 GB)48GB91.1 tok/sComputingForGeeks, Ollama

All figures are community-reported. File sizes come from the Ollama registry.

Two testers disagree by 1.8x on the same file size. DeepSeek-R1 32B and Qwen2.5 32B at Q4_K_M are both 19.85 GB in Ollama. Hostkey measured 50.77 tok/s on the 72GB card. ComputingForGeeks measured 27.8 tok/s on the 48GB card. The two variants have the same bandwidth, so memory size does not explain the gap. The runtime and settings do. Benchmark your own stack before you blame the card.

Which variant should you buy?

Buy the 72GB if you want gpt-oss 120B on one card, or two mid-size models loaded at once. It costs $1,500 more at Central Computer for 50% more memory. At Micro Center, search results showed the PNY 72GB at $9,999.99, in-store pickup only. There is no catalog Amazon link for the 72GB card on this site. Buy it from a workstation retailer and confirm “72GB” on the listing.

Buy the 48GB if your models are 35B-class or smaller. It runs them at the same speed as the 72GB. The PNY RTX PRO 5000 Blackwell 48GB link goes to the 48GB card. Confirm the memory size on the listing before you pay, because both variants share a name.

Price check, as of September 2026:

ListingPriceStatusSource
PNY 48GB, Central Computer$7,499.99In stockRetailer page, read 2026-09-27
NVIDIA 72GB OEM, Central Computer$8,999.99Unavailable onlineRetailer page, read 2026-09-27
PNY 48GB, Newegg$9,199.00Third-party sellerRetailer page, read 2026-09-27
PNY 48GB, Micro Center$6,999.99In-store pickupSearch-result snippet, 2026-09-27
PNY 72GB, Micro Center$9,999.99In-store pickup onlySearch-result snippet, 2026-09-27

Micro Center blocked our direct fetch. We read its prices from search-result snippets on 2026-09-27. Confirm them in store. Pangoly’s history for the PNY 48GB card ranges from $6,049 (2026-07-17) to $9,199 (2026-08-12). Prices on this card move by thousands of dollars within weeks.

MIG is a real option on this card. It splits into two isolated instances, 2 x 36GB on the 72GB card. Each half can hold Qwen3.8 27B at Q6_K (20.47 GiB) with 128K of context (8 GiB), by our arithmetic. See MIG for local LLMs for the setup.

RTX PRO 5000 vs two RTX 5090s, the RTX PRO 6000 and a used A6000

Prices as of September 2026.

SetupVRAMBandwidthPowerPrice$/GBgpt-oss 120B + 128K?
RTX PRO 5000 48GB48GB1,344 GB/s300W$7,499.99$156No
RTX PRO 5000 72GB72GB1,344 GB/s300W$8,999.99$125Yes, ~8 GiB spare
2x RTX 509064GB (2 x 32)1,792 GB/s each1,150W~$8,600 (2 x $4,299.99 deal)$134Tight
RTX PRO 6000 Workstation96GB1,792 GB/s600W$15,999+$167Yes, ~30 GiB spare
RTX PRO 6000 Max-Q96GB1,792 GB/s300W$17,999.99 (Newegg)$187Yes
Used RTX A600048GB GDDR6768 GB/s300Wfrom $4,400 (eBay)$92No

Price sources: Central Computer (2026-09-27); Walmart’s MSI RTX 5090 Gaming Trio GeForce Week deal at $4,299.99, which sold out; Thunder Compute’s RTX PRO 6000 tracker (2026-09-18: Newegg and B&H $15,999+); Newegg’s own Max-Q listing (2026-09-27); GPUDojo’s used A6000 page (2026-09-27).

Two RTX 5090s: faster, hotter, and tighter. Each 5090 has 1,792 GB/s, but a model split across two cards does not decode twice as fast. On a 70B Q4 model, measured decode was close. Database Mart measured 27.03 tok/s on two 5090s with Ollama. Hostkey measured 25.08 tok/s on the PRO 5000 72GB. The pair draws 1,150W against 300W. gpt-oss 120B at full context leaves only a few GiB across two cards for cache and two sets of buffers. The $4,299.99 price was a deal that sold out. Marketplace listings run above $6,000 per card. If you still want the pair, the GIGABYTE RTX 5090 32GB is the catalog pick. Check the seller before you pay. See what PSU for a local AI rig for a 1,150W build.

RTX PRO 6000: 24GB more for about $7,000 more. It holds gpt-oss 120B with about 30 GiB spare, and it has 33% more bandwidth. Hostkey measured it at 30.19 tok/s on DeepSeek-R1 70B and 58.73 tok/s on the 32B. The PRO 5000 72GB did 25.08 and 50.77 on the same tests. The NVIDIA RTX PRO 6000 Blackwell 96GB link is the 600W Workstation Edition.

The “$8,299 RTX PRO 6000” is not a current price. That Newegg Max-Q deal was posted on Slickdeals on 2025-09-28 and has expired. The TechRadar “massive price cut” stories are 2025 preorder news. On 2026-09-27, Newegg sold the Max-Q itself at $17,999.99.

Used RTX A6000: the cheapest 48GB, at 57% of the bandwidth. It has 768 GB/s against the PRO 5000’s 1,344 GB/s. At $92/GB it wins on price, if your models fit in 48GB. The PNY RTX A6000 48GB is the catalog pick. Used A6000 prices have risen: GPUDojo showed “used from $4,400” on 2026-09-27.

Caveats

  1. Hostkey’s gpt-oss 120B figure has no stated runtime or context length. It is the only measurement we found for this model on this card.
  2. The 72GB usable-memory figure is our estimate. We scaled the 48GB card’s 48,935 MiB. Check nvidia-smi on your card.
  3. Retail stock is thin. The 72GB listing at Central Computer was unavailable online. The Micro Center 72GB is in-store pickup only.
  4. Our fit budgets assume the runtime honors gpt-oss’s 128-token sliding window. A runtime that caches all 36 layers at full length uses 9.0 GiB, not 4.5 GiB. That still fits the 72GB card.
  5. The estimates rest on one MoE measurement and two dense ones. Treat them as a range, not a promise.

FAQ

What is the best local LLM for the RTX PRO 5000?

On the 72GB card, gpt-oss 120B. Its MXFP4 GGUF is 59.03 GiB and its full 131,072-token KV cache is about 4.5 GiB, so the whole job is about 63.5 GiB. Hostkey measured 156.52 tok/s for gpt-oss 120B on the 72GB card (community-reported, runtime not stated). On the 48GB card, the best pick is Qwen3.6 35B-A3B at Q8_0 (34.37 GiB), which fits with its full 262K context.

Can the RTX PRO 5000 48GB run gpt-oss 120B?

Not on the GPU alone. The 48GB card reports 48,935 MiB (47.8 GiB) in nvidia-smi. The gpt-oss 120B MXFP4 file is 59.03 GiB, so the weights miss by about 11 GiB before any context. You can run it with some expert layers in system RAM, but decode speed then depends on your CPU memory bandwidth.

Should I buy the RTX PRO 5000 48GB or 72GB?

Buy the 72GB if you want gpt-oss 120B or two 27-35B models resident at once. As of September 2026, Central Computer listed the 48GB at $7,499.99 (in stock) and the 72GB at $8,999.99 (unavailable online). That is $1,500 more for 24GB more, and $125 per GB against $156 per GB. Both cards have the same 1,344 GB/s bandwidth, so a model that fits in 48GB runs at the same speed on either.

RTX PRO 5000 72GB or two RTX 5090s?

Two RTX 5090s have 64GB, 1,792 GB/s each, and draw 575W each. At the $4,299.99 Walmart deal price (sold out in September 2026), a pair costs about $8,600. The PRO 5000 72GB has 8GB more memory on one card at 300W. On a 70B Q4 model, measured decode was close: 27.03 tok/s on two 5090s (Database Mart, Ollama) and 25.08 tok/s on the PRO 5000 72GB (Hostkey). gpt-oss 120B at full context is tight on 64GB and comfortable on 72GB.

How much VRAM and bandwidth does the RTX PRO 5000 have?

48GB or 72GB of GDDR7 with ECC, both at 1,344 GB/s on a 384-bit bus, per NVIDIA and PNY. Both variants have 14,080 CUDA cores, a 300W maximum power draw, a dual-slot cooler, PCIe 5.0, and MIG support for up to two isolated instances. The bandwidth is 75% of an RTX 5090's or RTX PRO 6000's 1,792 GB/s.

Sources

Before you order parts, check the tested hardware list for current prices by tier.

See Also

Need OpenClaw fixed live?

Remote rescue sessions for gateway, auth, tunnel, VPS, and model access problems.

See Rescue Session

Read next

Is 96GB of VRAM Enough for Local AI in 2026?
96GB is the first tier where a dense 70B runs at its full 128K window: 42.5GB of Q4 weights plus exactly 40 GiB of FP16 KV cache is 82.5GB, and it fits. What 96GB unlocks, what it still cannot hold, and what the one card that has it costs in 2026.
Best 48GB VRAM Setup for Local LLMs
Four routes to 48GB of VRAM: two used RTX 3090s ($2,000-2,600), a used RTX A6000 ($2,600-3,800), an RTX PRO 5000 Blackwell 48GB ($5,600-6,250), or skip to 96GB. Which one to buy, and the 700W tax nobody prices in.
Two Used RTX 3090s or One RTX 5090? 48GB Slow vs 32GB Fast
Dual used RTX 3090s cost $2,000-2,600 for 48GB of VRAM. One RTX 5090 costs $4,300-5,000 for 32GB. The 2026 price spike flipped this comparison: the dual build is now half the price AND holds a 70B. Here is the honest tradeoff, including the 700W problem.
Best Local LLM for 64GB VRAM (2026): Laguna S 2.1 Wins
Best local LLM for 64GB VRAM: Laguna S 2.1 UD-IQ4_XS (53.6 GiB) split over two RTX 5090s, Qwen3.8 27B Q8 on one card. Why gpt-oss 120B does not fit.