Is the RTX 4090 Worth $2,600 for Local LLMs? (2026)
The RTX 4090 launched at $1,599 in October 2022, and NVIDIA stopped making it in late 2024. As of September 2026 a used one sells for about $2,600 on eBay, and new stock costs $3,200 or more. That is more per GB/s of bandwidth than a new RTX 5090 at its $4,299 floor. This page gives the verdict at today's price for each type of buyer, shows what the same money buys instead, and works out how many rental hours the card has to replace.
Bottom Line
A used RTX 4090 costs about $2,600 as of September 2026. That is 1.6x its $1,599 launch price from 2022. Whether it is worth that depends on who you are.
| Buyer | Worth ~$2,600? | Why |
|---|---|---|
| You run long-context agents or coding on 24GB, need CUDA, and cannot spend $4,299 | Yes | It processes prompts 2.3-2.7x faster than a 3090 in community llama.cpp tests. Long prompts wait less. |
| You already own one | Yes, keep it | It resells above its launch price. Nothing new under $4,000 is faster on CUDA. |
| You chat with 20-35B models | No | A used RTX 3090 ($1,463) has the same 24GB and generates only about 15% slower, for 56% of the price. |
| You want 70B dense or 100B+ MoE models | No | 24GB cannot hold Llama 3.3 70B Q4_K_M (42.5GB). Two R9700s at MSRP ($2,598) give 64GB for the same money. |
| You can stretch to $4,299 | No | The RTX 5090 costs less per GB/s ($2.40 against $2.57) and adds 8GB. |
| You use local AI a few hours a week | No | RunPod rents a 4090 for $0.34-0.74 per hour. The card needs about 3,930 hours of use to beat the higher rate. |
The short version: at 2026 prices the 4090 has the most expensive memory bandwidth of the four 24-32GB cards we price. It costs more per GB/s than a used 3090, an R9700, and even a new RTX 5090 at its $4,299 floor. It sits between those cards on speed, but not on value. Buy it only for prefill speed on 24GB, at a budget that stops short of a 5090.
The price as of September 2026
NVIDIA told board partners to stop RTX 4090 production by October 2024 (TweakTown, 2024-10-06, citing Board Channels). New stock is now leftover inventory. The prices we found tonight:
| Measure | Price | Date | Source |
|---|---|---|---|
| Launch MSRP (Founders Edition) | $1,599 | 12 Oct 2022 | Wikipedia, GeForce RTX 40 series |
| Used, latest eBay sold reading | $2,510 | 2026-09-19 | GetPCParts |
| Used, 30-day average of eBay sold readings | $2,589 | to 2026-09-19 | GetPCParts (5 readings) |
| Used, highest reading this year | $3,010 | 2026-09-12 | GetPCParts |
| Used, eBay US median | $3,000 | 2026-09-25 | ResalePrices (17 listings, 90-day range $2,497-3,075, +20% in 30 days) |
| New, Amazon | $3,200 | read 2026-09-25 | GPUDojo |
| New, Newegg | $4,499 | read 2026-09-25 | GPUDojo |
| August 2026 used range (our earlier read) | $2,270-2,600 | August 2026 | ResalePrices / BestValueGPU |
We use $2,589 (the GetPCParts 30-day sold average) as the working price on this page. ResalePrices shows $3,000, but from only 17 listings. GPUDojo also lists an eBay “new” price of $2,073. We treat that as one marketplace listing, not a market price.
The 4090 sells above its launch price, and the gap is growing. At $2,589 it costs 162% of its $1,599 MSRP. A used RTX 3090 sells at 97% of its launch price. The cause is the 2026 memory shortage, plus a card that no one makes any more. Prices can also fall. A card at 1.6x its own MSRP has room to drop if new-card prices ease.
If you buy, compare the listing against the $2,500-2,600 sold range first. The catalog pick is the RTX 4090 24GB. Check the listed price: it moves between $3,200 and $4,500 new.
What $2,600 buys instead
All prices as of September 2026. Dollars per GB and per GB/s are our arithmetic, with the same method as our RTX 5090 and RTX 3090 verdicts. Bandwidth comes from vendor spec sheets.
| Option | Price | Memory | Bandwidth | $ per GB | $ per GB/s | Holds Llama 3.3 70B Q4_K_M (42.5GB)? |
|---|---|---|---|---|---|---|
| Used RTX 4090 | $2,589 | 24GB GDDR6X | 1,008 GB/s | $108 | $2.57 | No |
| Used RTX 3090 | $1,463 avg | 24GB GDDR6X | 936 GB/s | $61 | $1.56 | No |
| 2x used RTX 3090 | $2,926 | 48GB | 936 GB/s each | $61 | n/a | Yes |
| Radeon AI PRO R9700 | $1,299 MSRP; $1,599 street avg | 32GB GDDR6 | 640 GB/s | $41 at MSRP; $50 street | $2.03 at MSRP; $2.50 street | No |
| 2x R9700 | $2,598 at MSRP | 64GB | 640 GB/s each | $41 | n/a | Yes |
| RTX 5060 Ti 16GB (new) | $789 (B&H, 2026-09-24) | 16GB GDDR7 | 448 GB/s | $49 | $1.76 | No |
| RTX 5090 (new) | $4,299 floor (2026-09-22) | 32GB GDDR7 | 1,792 GB/s | $134 | $2.40 | No |
| RunPod RTX 4090 rental | $0.34-0.74/hr | 24GB | 1,008 GB/s | n/a | n/a | No |
Three things stand out.
- On capacity, the 4090 is the second-worst buy. At $108 per GB it costs 1.8x a used 3090 per GB. Only the 5090 costs more, at $134.
- On bandwidth, the 4090 is the worst buy. At $2.57 per GB/s it costs more than the 5090 ($2.40), the R9700 at street price ($2.50) and the 3090 ($1.56). Token generation speed follows bandwidth.
- The exact price of one used 4090 buys two R9700s at MSRP. That is 64GB against 24GB. It holds a 70B model at Q4_K_M. Stock is the catch: Newegg showed the R9700 out of stock on 2026-09-24.
The three alternatives, in numbers
Used RTX 3090. Same 24GB, same model fit. Its 936 GB/s is 93% of the 4090’s 1,008 GB/s. On the llama.cpp CUDA scoreboard it generated 158-162 tok/s on Llama 2 7B Q4_0, against 186-189 for the 4090. ResalePrices shows a $1,463 average across 291 eBay US listings on 2026-09-24. A used RTX 3090 24GB costs 56% of a 4090. Pick a model with which used RTX 3090 to buy.
RTX 5090. It costs 1.66x the used 4090 ($4,299 against $2,589). It has 1.78x the bandwidth and 8GB more memory. In community tests it generated 1.3-1.56x faster (details below). NVIDIA still sells new 5090s with a warranty. Used 4090s have none from the seller. If you can reach $4,299 from a first-party seller, the RTX 5090 is the better use of money. The verdict is in is the RTX 5090 worth it.
Radeon AI PRO R9700 32GB. It has 8GB more than the 4090 at 50-62% of the price. It has 63% of the 4090’s bandwidth (640 against 1,008 GB/s), so expect slower generation on a model that fits both. That is a bandwidth-ratio estimate, not a measurement. It runs on ROCm, not CUDA. Buy the Radeon AI PRO R9700 32GB near its $1,299 MSRP. The comparison against the 3090 is in R9700 vs RTX 3090.
RTX 4090 LLM performance on models that fit in 24GB
All numbers below are community-reported llama-bench results from the llama.cpp GitHub discussions. They are not our measurements.
| Model | Test | RTX 3090 | RTX 4090 | RTX 5090 | Source |
|---|---|---|---|---|---|
| Llama 2 7B Q4_0 | Generation (tg128), FA off | 158.2 tok/s | 186.2 tok/s | 290.0 tok/s | Discussion #15013 |
| Llama 2 7B Q4_0 | Prompt (pp512), FA off | 5,175 tok/s | 11,993 tok/s | 14,073 tok/s | Discussion #15013 |
| Llama 2 7B Q4_0 | Generation, FA on | 161.9 tok/s | 189.0 tok/s | not listed | Discussion #15013 |
| Llama 2 7B Q4_0 | Prompt, FA on | 5,560 tok/s | 14,771 tok/s | not listed | Discussion #15013 |
| Qwen3 30B-A3B Q4_K | Generation, CUDA graph opt on | not tested | 271.0 tok/s | 352.1 tok/s | Discussion #17621 (Nov 2025) |
| gpt-oss 20B MXFP4 | Generation, CUDA graph opt on | not tested | 272.0 tok/s | 419.1 tok/s | Discussion #17621 (Nov 2025) |
The 3090 and 4090 rows on the scoreboard come from different llama.cpp builds. Read them as a guide, not an exact ratio.
What the numbers say:
- Generation: the 4090 is only 17-18% faster than a 3090. 186.2 / 158.2 = 1.18. 189.0 / 161.9 = 1.17. The bandwidth gap is 8% (1,008 / 936).
- Prompt processing: the 4090 is 2.3-2.7x faster than a 3090. 11,993 / 5,175 = 2.3. 14,771 / 5,560 = 2.7. This is where the extra money goes.
- The 5090 is 1.3-1.56x faster than the 4090 at generation. 290.0 / 186.2 = 1.56. 352.1 / 271.0 = 1.30. 419.1 / 272.0 = 1.54.
Now divide the price by the Llama 2 7B generation speed (FA off). Our arithmetic:
- Used RTX 3090: $1,463 / 158.2 = $9.25 per tok/s
- Used RTX 4090: $2,589 / 186.2 = $13.90 per tok/s
- RTX 5090: $4,299 / 290.0 = $14.82 per tok/s
The 4090 costs 50% more per unit of generation speed than the 3090. It costs only 6% less than the 5090.
What 24GB runs
File sizes from the Hugging Face API, as listed on our RTX 3090 verdict (read 2026-09-24).
| Model | Quant | File size | Fits 24GB? | 4090 generation |
|---|---|---|---|---|
| gpt-oss 20B | native MXFP4 | 13.8GB | Yes | 272 tok/s (community, #17621) |
| Qwen3.6 27B | Q4_K_M | 16.8GB | Yes | ~48 tok/s (our estimate, below) |
| Qwen3.6 35B-A3B | MXFP4_MOE | 21.7GB | Yes, with little context room | not measured |
| Llama 3.3 70B | Q4_K_M | 42.5GB | No | n/a |
| gpt-oss 120B | MXFP4 | 65.4GB | No | n/a |
The Qwen3.6 27B figure is an estimate. A dense model reads its full weights for each token. 1,008 GB/s / 16.8GB = 60 tok/s as a ceiling. The #17621 author reached about 80% of theoretical bandwidth on a 5090. 60 x 0.8 = 48 tok/s. The model picks for this card are in best local LLM for the RTX 4090. For 70B, read can an RTX 4090 run a 70B model.
Rent vs buy
RunPod lists the RTX 4090 at $0.74 per hour on Secure Cloud and $0.34 on Community Cloud (runpod.io/pricing, read 2026-09-22). Owning the card also costs power. NVIDIA rates the 4090 at 450W total graphics power. At $0.18/kWh (US average, as of August 2026), that is about $0.08 per hour under load.
Our arithmetic, for $2,589:
| Rental rate | Net saving per hour of ownership | Break-even hours | At 4 h/day | At 8 h/day | 24/7 |
|---|---|---|---|---|---|
| $0.74 (Secure) | $0.66 | ~3,930 | 2.7 years | 1.3 years | 5.4 months |
| $0.34 (Community) | $0.26 | ~10,000 | 6.8 years | 3.4 years | 13.7 months |
The formula is $2,589 / ($0.74 - $0.081) = 3,929 hours. At the Community rate it is $2,589 / ($0.34 - $0.081) = 9,996 hours.
Two things favor buying. A card you own keeps resale value, and today’s resale value is above launch price. Your prompts also never leave your machine. One thing favors renting: the 4090 is nearly four years old and has no seller warranty. If you cannot name the job that fills 4 hours a day, rent for a month and log your hours first.
Is the RTX 4090 worth it over a used RTX 3090?
For chat, no. Both cards hold the same 24GB and run the same models. The 4090 generates 17-18% faster and costs 1.77x as much ($2,589 / $1,463).
For agents and coding, sometimes. Agent loops resend 20-30K tokens of context on every call. The 4090 processes prompts 2.3-2.7x faster than a 3090 on the scoreboard test. If your work waits on prompt processing, the 4090 pays for that wait. The full head-to-head is in RTX 3090 vs RTX 4090.
Is the RTX 4090 worth it over an RTX 5090?
Only if $4,299 is outside your budget. The 5090 costs 1.66x the used 4090. It gives 1.78x the bandwidth, 1.33x the memory (32GB against 24GB), and a new-card warranty. Per GB/s it is the cheaper card. The three-way comparison is in RTX 5090 vs 4090 vs used 3090.
Power and fit
NVIDIA rates the 4090 at 450W and specifies an 850W minimum system power supply. A used card has no warranty from the seller. Stress-test it before the return window closes. Our used RTX 3090 checklist applies to the 4090 as well.
Sources
- GetPCParts: used RTX 4090 eBay sold prices, read 2026-09-25 (latest reading 2026-09-19)
- ResalePrices: used RTX 4090 eBay US market, updated 2026-09-25
- GPUDojo: RTX 4090 new prices by retailer, read 2026-09-25
- TweakTown: NVIDIA to cease RTX 4090 production by October 2024, 2024-10-06
- NVIDIA GeForce RTX 4090 specifications (24GB GDDR6X, 384-bit, 450W, 850W system power)
- Wikipedia: GeForce RTX 40 series (launch 12 Oct 2022, $1,599 MSRP, 1,008 GB/s)
- llama.cpp discussion #15013: performance of llama.cpp on NVIDIA CUDA (community-reported)
- llama.cpp discussion #17621: optimizing token generation in the CUDA backend, 2025-11-30 (community-reported)
- ResalePrices: used RTX 3090 eBay US market, read 2026-09-24
- Tom’s Hardware: MSI RTX 5090 for $4,299 at Walmart, 2026-09-22
- AMD Radeon AI PRO R9700 specifications (32GB, 640 GB/s)
- runaihome: R9700 street price tracking, September 2026
- gpuprix: RTX 5060 Ti 16GB US price, read 2026-09-24
- RunPod GPU pricing, read 2026-09-22
- Model file sizes from the Hugging Face API: unsloth/gpt-oss-20b-GGUF, unsloth/Qwen3.6-27B-GGUF, unsloth/Qwen3.6-35B-A3B-GGUF, bartowski/Llama-3.3-70B-Instruct-GGUF, unsloth/gpt-oss-120b-GGUF
See Also
- Best Local LLM for the RTX 4090: what to run on 24GB once you own one
- Is the RTX 5090 Worth It?: the 32GB flagship at $4,299
- Is a Used RTX 3090 Worth It?: the same 24GB at $1,463
- RTX 5090 vs 4090 vs Used 3090: the three-card buying rule
- RTX 3090 vs RTX 4090 for Local LLMs: same VRAM, different speed
- Which Used RTX 3090 Should You Buy?: pick on cooling and slot width
- R9700 vs RTX 3090 for Local LLMs: 32GB AMD against 24GB CUDA
- Can an RTX 4090 Run a 70B Model?: the limit of one card
Need OpenClaw fixed live?
Remote rescue sessions for gateway, auth, tunnel, VPS, and model access problems.
See Rescue Session