Best Local LLM for RTX PRO 5000 (2026): gpt-oss 120B Wins
The RTX PRO 5000 Blackwell ships in two sizes: 48GB and 72GB of GDDR7, both at 1,344 GB/s. The 72GB card is the smallest single NVIDIA card that holds gpt-oss 120B with its full 128K context. The 48GB card misses it by 11 GiB, so its best model is Qwen3.6 35B-A3B at Q8_0. This page lists what fits on each variant, the measured speeds, and what two RTX 5090s or an RTX PRO 6000 give you for the same money.
Shopping? See the tested hardware list with every verified pick by tier.
Bottom Line
- Best model on the 72GB card: gpt-oss 120B. The MXFP4 file is 59.03 GiB. Its full 128K context costs about 4.5 GiB. The whole job fits on one card with room left.
- Measured speed: Hostkey reports 156.52 tok/s for gpt-oss 120B on the 72GB card. This is community-reported, and the runtime is not stated.
- Best model on the 48GB card: Qwen3.6 35B-A3B at Q8_0. The file is 34.37 GiB. It fits with the full 262K context. Only 3B parameters are active per token.
- The 48GB card cannot hold gpt-oss 120B. It reports 47.8 GiB usable. The weights alone are 11.2 GiB too large.
- Both variants have the same speed. They share 1,344 GB/s, 14,080 CUDA cores and 300W. The 72GB adds capacity only.
- Price, as of September 2026: 48GB at $7,499.99 and 72GB at $8,999.99 at Central Computer. The 72GB was listed as unavailable online.
The 72GB card is cheaper per gigabyte than the 48GB card. At the one US retailer that lists both, it costs $125/GB against $156/GB. It also undercuts the RTX PRO 6000 ($167/GB) and two RTX 5090s at their best deal price ($134/GB).
RTX PRO 5000 VRAM and bandwidth
| Specification | RTX PRO 5000 48GB | RTX PRO 5000 72GB |
|---|---|---|
| Memory | 48GB GDDR7 ECC | 72GB GDDR7 ECC |
| Bus width | 384-bit | 384-bit |
| Bandwidth | 1,344 GB/s | 1,344 GB/s |
| CUDA cores | 14,080 | 14,080 |
| Max power | 300W | 300W |
| Cooler | Dual slot, active | Dual slot, active |
| MIG | Up to 2 instances | Up to 2 x 36GB |
| Usable in nvidia-smi | 48,935 MiB (47.8 GiB), measured | ~71.7 GiB (our estimate) |
Specifications from NVIDIA’s RTX PRO 5000 page and PNY’s 72GB product page, read 2026-09-27. The 48GB nvidia-smi figure comes from a ComputingForGeeks test (2026-09-25). We found no nvidia-smi reading for the 72GB card. Our 71.7 GiB figure scales the 48GB reading by 1.5.
Decode speed follows bandwidth. For each token, a dense model reads all its weights. A MoE model reads only its active experts. So on a 1,344 GB/s card, MoE models are the fast choice.
Best Local LLMs for the RTX PRO 5000 72GB
1. gpt-oss 120B (MXFP4): the winner
gpt-oss 120B has 117B parameters with 5.1B active per token, per OpenAI’s model card. The native MXFP4 GGUF (ggml-org) is 63,387,346,208 bytes, which is 59.03 GiB.
Its KV cache is small. Half of its 36 layers use a 128-token sliding window. The other 18 layers cost 36 KiB per token at FP16. So the full 131,072-token window costs about 4.5 GiB. See gpt-oss 120B VRAM at full context for the arithmetic.
| Item | Size |
|---|---|
| Weights, MXFP4 GGUF | 59.03 GiB |
| KV cache, FP16, full 131K | 4.5 GiB |
| Total | 63.5 GiB |
| Left on the 72GB card | ~8 GiB (estimate) for buffers |
./build/bin/llama-server -hf ggml-org/gpt-oss-120b-GGUF -ngl 99 -c 131072 --jinja
Measured: 156.52 tok/s (Hostkey, 72GB card, 2026-09-08). Hostkey does not name its runtime or context length. Treat the number as one data point.
The line other reviews miss: the context is not what blocks the 48GB card. The full 128K cache is 4.5 GiB. The 48GB card is 11.2 GiB short on the weights alone. No context setting fixes that.
2. Two models resident at once
The 72GB card can hold Qwen3.6 35B-A3B at Q8_0 (34.37 GiB) and Qwen3.8 27B at Q8_0 (27.05 GiB) together. That is 61.4 GiB of weights. With 64K of context each, the caches add 1.25 GiB and 4.0 GiB. The total is about 66.7 GiB. This fit is tight. It is our arithmetic, not a measurement.
This is the agent setup: a fast MoE model for tool calls and a dense 27B for careful code.
3. Qwen3.8 27B at BF16: an unquantized 27B
The BF16 file is 50.90 GiB. With a 128K context (8 GiB), it uses about 58.9 GiB. You get the model with no quantization loss. The cost is speed: a dense 55 GB read per token.
Best Local LLMs for the RTX PRO 5000 48GB
1. Qwen3.6 35B-A3B (Q8_0): the winner
Qwen3.6 35B-A3B has 35B total parameters and 3B active, with 256 experts. Only 10 of its 40 layers use full attention. From the config (2 KV heads, head_dim 256), that is 20 KiB per token at FP16. Its native 262,144-token context costs 5.0 GiB.
At Q8_0 (34.37 GiB) plus the full 262K cache, the total is 39.4 GiB. That leaves about 8 GiB on the 48GB card. You run the model at 8-bit with its whole window.
2. Qwen3.8 27B (Q8_0): best dense model
Qwen3.8 27B has 64 layers, and 16 use full attention (4 KV heads, head_dim 256). That is 64 KiB per token, so 262K costs 16 GiB. The Q8_0 file is 27.05 GiB. Model plus the full window is 43.1 GiB. It fits on the 48GB card, with about 4.7 GiB left.
3. gpt-oss 20B: the lightweight agent model
gpt-oss 20B has 21B parameters and 3.6B active. The unsloth F16 file, with MXFP4 experts, is 12.85 GiB. Its full 131K context costs 3.0 GiB. It uses under a third of the card.
What about gpt-oss 120B on 48GB?
It does not fit on the GPU. You can keep some expert layers in system RAM with llama.cpp’s --n-cpu-moe option. Speed then depends on your CPU memory. We found no measurement on this card, so we give no number.
What fits on each variant
File sizes from the Hugging Face API, read 2026-09-27. Context budgets are our arithmetic from each model’s config.json at FP16 KV. We keep 2 GiB free for compute buffers.
| Model | Quant | File | 48GB card | 72GB card | Decode (label) |
|---|---|---|---|---|---|
| gpt-oss 120B | MXFP4 | 59.03 GiB | No | Yes, full 131K | 156.52 tok/s, measured (Hostkey) |
| Qwen3.6 35B-A3B | Q8_0 | 34.37 GiB | Yes, full 262K | Yes, full 262K | ~140 tok/s, estimate |
| Qwen3.8 27B | Q8_0 | 27.05 GiB | Yes, full 262K | Yes, full 262K | 19-37 tok/s, estimate |
| Qwen3.8 27B | BF16 | 50.90 GiB | No | Yes, 128K | 10-19 tok/s, estimate |
| gpt-oss 20B | F16 (MXFP4 experts) | 12.85 GiB | Yes, full 131K | Yes, full 131K | ~185 tok/s, estimate |
| Qwen3.6 35B-A3B + Qwen3.8 27B | Q8_0 + Q8_0 | 61.42 GiB | No | Yes, 64K each (tight) | n/a |
| Qwen3.8 Flash-Next | UD-IQ1_S (smallest) | 67.56 GiB | No | Loads, almost no context | Not recommended |
| DeepSeek V4 Flash 0731 | UD-IQ1_S (smallest) | 76.87 GiB | No | No | n/a |
| GLM-5.3 Flash | UD-IQ1_S (smallest) | 86.69 GiB | No | No | n/a |
How we derived the estimates. We read the ceiling as 1,344 GB/s divided by the bytes read per token. Then we apply the efficiency this card reached in published tests:
- MoE models, about 33% of the ceiling. Hostkey’s 156.52 tok/s on gpt-oss 120B is about a third of its ceiling. That ceiling is 1,344 ÷ ~2.85 GB of active weights per token, or ~470 tok/s.
- Qwen3.6 35B-A3B Q8_0: 36.9 GB x 3/35 = 3.16 GB per token. 1,344 ÷ 3.16 = 425 tok/s ceiling. At 33%, about 140 tok/s.
- gpt-oss 20B: 13.8 GB x 3.6/21 = 2.37 GB per token. Ceiling 568 tok/s. At 33%, about 185 tok/s.
- Dense models, 41-79% of the ceiling. That is the range between two published tests (see below).
- Qwen3.8 27B Q8_0: 29.05 GB per token. Ceiling 46.3 tok/s. At 41-79%, 19-37 tok/s.
These are estimates, not measurements. Your runtime, context length and settings move them.
Measured results on this chip
| Model (file size) | Card | Decode | Source |
|---|---|---|---|
| gpt-oss 120B | 72GB | 156.52 tok/s | Hostkey, 2026-09-08 |
| Qwen3-Coder-Next q4_K_M | 72GB | 192.13 tok/s | Hostkey |
| DeepSeek-R1 70B (42.52 GB) | 72GB | 25.08 tok/s | Hostkey |
| DeepSeek-R1 32B (19.85 GB) | 72GB | 50.77 tok/s | Hostkey |
| Qwen2.5 32B Q4_K_M (19.85 GB) | 48GB | 27.8 tok/s | ComputingForGeeks, Ollama, 2026-09-25 |
| Qwen2.5 14B Q4_K_M | 48GB | 54.7 tok/s | ComputingForGeeks, Ollama |
| Llama 3.1 8B Q4_K_M (4.92 GB) | 48GB | 91.1 tok/s | ComputingForGeeks, Ollama |
All figures are community-reported. File sizes come from the Ollama registry.
Two testers disagree by 1.8x on the same file size. DeepSeek-R1 32B and Qwen2.5 32B at Q4_K_M are both 19.85 GB in Ollama. Hostkey measured 50.77 tok/s on the 72GB card. ComputingForGeeks measured 27.8 tok/s on the 48GB card. The two variants have the same bandwidth, so memory size does not explain the gap. The runtime and settings do. Benchmark your own stack before you blame the card.
Which variant should you buy?
Buy the 72GB if you want gpt-oss 120B on one card, or two mid-size models loaded at once. It costs $1,500 more at Central Computer for 50% more memory. At Micro Center, search results showed the PNY 72GB at $9,999.99, in-store pickup only. There is no catalog Amazon link for the 72GB card on this site. Buy it from a workstation retailer and confirm “72GB” on the listing.
Buy the 48GB if your models are 35B-class or smaller. It runs them at the same speed as the 72GB. The PNY RTX PRO 5000 Blackwell 48GB link goes to the 48GB card. Confirm the memory size on the listing before you pay, because both variants share a name.
Price check, as of September 2026:
| Listing | Price | Status | Source |
|---|---|---|---|
| PNY 48GB, Central Computer | $7,499.99 | In stock | Retailer page, read 2026-09-27 |
| NVIDIA 72GB OEM, Central Computer | $8,999.99 | Unavailable online | Retailer page, read 2026-09-27 |
| PNY 48GB, Newegg | $9,199.00 | Third-party seller | Retailer page, read 2026-09-27 |
| PNY 48GB, Micro Center | $6,999.99 | In-store pickup | Search-result snippet, 2026-09-27 |
| PNY 72GB, Micro Center | $9,999.99 | In-store pickup only | Search-result snippet, 2026-09-27 |
Micro Center blocked our direct fetch. We read its prices from search-result snippets on 2026-09-27. Confirm them in store. Pangoly’s history for the PNY 48GB card ranges from $6,049 (2026-07-17) to $9,199 (2026-08-12). Prices on this card move by thousands of dollars within weeks.
MIG is a real option on this card. It splits into two isolated instances, 2 x 36GB on the 72GB card. Each half can hold Qwen3.8 27B at Q6_K (20.47 GiB) with 128K of context (8 GiB), by our arithmetic. See MIG for local LLMs for the setup.
RTX PRO 5000 vs two RTX 5090s, the RTX PRO 6000 and a used A6000
Prices as of September 2026.
| Setup | VRAM | Bandwidth | Power | Price | $/GB | gpt-oss 120B + 128K? |
|---|---|---|---|---|---|---|
| RTX PRO 5000 48GB | 48GB | 1,344 GB/s | 300W | $7,499.99 | $156 | No |
| RTX PRO 5000 72GB | 72GB | 1,344 GB/s | 300W | $8,999.99 | $125 | Yes, ~8 GiB spare |
| 2x RTX 5090 | 64GB (2 x 32) | 1,792 GB/s each | 1,150W | ~$8,600 (2 x $4,299.99 deal) | $134 | Tight |
| RTX PRO 6000 Workstation | 96GB | 1,792 GB/s | 600W | $15,999+ | $167 | Yes, ~30 GiB spare |
| RTX PRO 6000 Max-Q | 96GB | 1,792 GB/s | 300W | $17,999.99 (Newegg) | $187 | Yes |
| Used RTX A6000 | 48GB GDDR6 | 768 GB/s | 300W | from $4,400 (eBay) | $92 | No |
Price sources: Central Computer (2026-09-27); Walmart’s MSI RTX 5090 Gaming Trio GeForce Week deal at $4,299.99, which sold out; Thunder Compute’s RTX PRO 6000 tracker (2026-09-18: Newegg and B&H $15,999+); Newegg’s own Max-Q listing (2026-09-27); GPUDojo’s used A6000 page (2026-09-27).
Two RTX 5090s: faster, hotter, and tighter. Each 5090 has 1,792 GB/s, but a model split across two cards does not decode twice as fast. On a 70B Q4 model, measured decode was close. Database Mart measured 27.03 tok/s on two 5090s with Ollama. Hostkey measured 25.08 tok/s on the PRO 5000 72GB. The pair draws 1,150W against 300W. gpt-oss 120B at full context leaves only a few GiB across two cards for cache and two sets of buffers. The $4,299.99 price was a deal that sold out. Marketplace listings run above $6,000 per card. If you still want the pair, the GIGABYTE RTX 5090 32GB is the catalog pick. Check the seller before you pay. See what PSU for a local AI rig for a 1,150W build.
RTX PRO 6000: 24GB more for about $7,000 more. It holds gpt-oss 120B with about 30 GiB spare, and it has 33% more bandwidth. Hostkey measured it at 30.19 tok/s on DeepSeek-R1 70B and 58.73 tok/s on the 32B. The PRO 5000 72GB did 25.08 and 50.77 on the same tests. The NVIDIA RTX PRO 6000 Blackwell 96GB link is the 600W Workstation Edition.
The “$8,299 RTX PRO 6000” is not a current price. That Newegg Max-Q deal was posted on Slickdeals on 2025-09-28 and has expired. The TechRadar “massive price cut” stories are 2025 preorder news. On 2026-09-27, Newegg sold the Max-Q itself at $17,999.99.
Used RTX A6000: the cheapest 48GB, at 57% of the bandwidth. It has 768 GB/s against the PRO 5000’s 1,344 GB/s. At $92/GB it wins on price, if your models fit in 48GB. The PNY RTX A6000 48GB is the catalog pick. Used A6000 prices have risen: GPUDojo showed “used from $4,400” on 2026-09-27.
Caveats
- Hostkey’s gpt-oss 120B figure has no stated runtime or context length. It is the only measurement we found for this model on this card.
- The 72GB usable-memory figure is our estimate. We scaled the 48GB card’s 48,935 MiB. Check
nvidia-smion your card. - Retail stock is thin. The 72GB listing at Central Computer was unavailable online. The Micro Center 72GB is in-store pickup only.
- Our fit budgets assume the runtime honors gpt-oss’s 128-token sliding window. A runtime that caches all 36 layers at full length uses 9.0 GiB, not 4.5 GiB. That still fits the 72GB card.
- The estimates rest on one MoE measurement and two dense ones. Treat them as a range, not a promise.
FAQ
What is the best local LLM for the RTX PRO 5000?
On the 72GB card, gpt-oss 120B. Its MXFP4 GGUF is 59.03 GiB and its full 131,072-token KV cache is about 4.5 GiB, so the whole job is about 63.5 GiB. Hostkey measured 156.52 tok/s for gpt-oss 120B on the 72GB card (community-reported, runtime not stated). On the 48GB card, the best pick is Qwen3.6 35B-A3B at Q8_0 (34.37 GiB), which fits with its full 262K context.
Can the RTX PRO 5000 48GB run gpt-oss 120B?
Not on the GPU alone. The 48GB card reports 48,935 MiB (47.8 GiB) in nvidia-smi. The gpt-oss 120B MXFP4 file is 59.03 GiB, so the weights miss by about 11 GiB before any context. You can run it with some expert layers in system RAM, but decode speed then depends on your CPU memory bandwidth.
Should I buy the RTX PRO 5000 48GB or 72GB?
Buy the 72GB if you want gpt-oss 120B or two 27-35B models resident at once. As of September 2026, Central Computer listed the 48GB at $7,499.99 (in stock) and the 72GB at $8,999.99 (unavailable online). That is $1,500 more for 24GB more, and $125 per GB against $156 per GB. Both cards have the same 1,344 GB/s bandwidth, so a model that fits in 48GB runs at the same speed on either.
RTX PRO 5000 72GB or two RTX 5090s?
Two RTX 5090s have 64GB, 1,792 GB/s each, and draw 575W each. At the $4,299.99 Walmart deal price (sold out in September 2026), a pair costs about $8,600. The PRO 5000 72GB has 8GB more memory on one card at 300W. On a 70B Q4 model, measured decode was close: 27.03 tok/s on two 5090s (Database Mart, Ollama) and 25.08 tok/s on the PRO 5000 72GB (Hostkey). gpt-oss 120B at full context is tight on 64GB and comfortable on 72GB.
How much VRAM and bandwidth does the RTX PRO 5000 have?
48GB or 72GB of GDDR7 with ECC, both at 1,344 GB/s on a 384-bit bus, per NVIDIA and PNY. Both variants have 14,080 CUDA cores, a 300W maximum power draw, a dual-slot cooler, PCIe 5.0, and MIG support for up to two isolated instances. The bandwidth is 75% of an RTX 5090's or RTX PRO 6000's 1,792 GB/s.
Sources
- NVIDIA RTX PRO 5000 Blackwell product page (48GB or 72GB GDDR7 ECC, 1,344 GB/s, 300W, dual slot, MIG up to 2 instances), read 2026-09-27
- PNY RTX PRO 5000 72GB product page (384-bit, 1,344 GB/s, 14,080 CUDA cores, MIG 2 x 36GB), read 2026-09-27
- NVIDIA blog: RTX PRO 5000 72GB generally available, 2025-12-18
- Hostkey: RTX PRO 5000 Blackwell 72GB tests, 2026-09-08
- ComputingForGeeks: RTX Pro 4000 vs RTX Pro 5000 local AI benchmarks (48GB card, Ollama, 48,935 MiB), 2026-09-25
- Database Mart: 2x RTX 5090 Ollama benchmark, updated 2026-08-25
- Central Computer: RTX PRO 5000 Blackwell listings (48GB $7,499.99 in stock; 72GB $8,999.99 unavailable online), read 2026-09-27
- Newegg: PNY RTX PRO 5000 48GB ($9,199.00, third-party seller), read 2026-09-27
- Micro Center: PNY RTX PRO 5000 72GB (price read from a search-result snippet, 2026-09-27)
- Pangoly: PNY RTX PRO 5000 48GB price history (read from a search-result snippet, 2026-09-27)
- Thunder Compute: RTX PRO 6000 pricing, 2026-09-18
- Newegg: NVIDIA RTX PRO 6000 Blackwell Max-Q ($17,999.99, sold by Newegg), read 2026-09-27
- Slickdeals: RTX PRO 6000 Max-Q $8,299.99 at Newegg, posted 2025-09-28, expired
- Tom’s Hardware: MSI RTX 5090 at $4,299 in Walmart’s GeForce Week and VideoCardz: the deal sold out
- PCGamesN: RTX 5090 listings over $6,000, official retailers out of stock, 2026-09-14
- GPUDojo: used RTX A6000 from $4,400, read 2026-09-27 and PNY RTX A6000 specifications (768 GB/s, 384-bit)
- Model file sizes from the Hugging Face API, read 2026-09-27: ggml-org/gpt-oss-120b-GGUF, unsloth/Qwen3.6-35B-A3B-GGUF, unsloth/Qwen3.8-27B-GGUF, unsloth/gpt-oss-20b-GGUF, unsloth/Qwen3.8-Flash-Next-GGUF, unsloth/DeepSeek-V4-Flash-0731-GGUF, unsloth/GLM-5.3-Flash-GGUF. Configs and parameter counts: openai/gpt-oss-120b, openai/gpt-oss-20b, Qwen/Qwen3.6-35B-A3B, Qwen/Qwen3.8-27B. Ollama file sizes from registry.ollama.ai.
Before you order parts, check the tested hardware list for current prices by tier.
See Also
- How Much VRAM Does gpt-oss 120B Need at Full 128K Context?: the 4.5 GiB cache arithmetic
- Best Local LLM for 96GB of VRAM: the RTX PRO 6000 tier above this card
- Best 48GB VRAM Setup for Local LLMs: the PRO 5000 48GB against dual 3090s and the A6000
- RTX PRO 6000 vs RTX 5090 for Local LLMs: capacity against price at the top end
- Best Local LLM for RTX 5090: the 32GB card at 1,792 GB/s
- Best Local LLM for RTX A6000: the used 48GB workstation card
- Best Local LLM for 64GB of VRAM: the dual-card tier
- RTX PRO 6000 Max-Q vs Workstation Edition: the 300W and 600W 96GB cards
- Is There a 48GB RTX 5090?: the 48GB routes on NVIDIA
Need OpenClaw fixed live?
Remote rescue sessions for gateway, auth, tunnel, VPS, and model access problems.
See Rescue Session