Two Used RTX 3090s or One RTX 5090? 48GB Slow vs 32GB Fast (August 2026)
Two used RTX 3090s give you 48GB of VRAM. One RTX 5090 gives you 32GB that is roughly three times faster per gigabyte moved. In 2025 these two options cost about the same, which made the choice interesting. As of August 2026 they do not: the dual-3090 build runs $2,000-2,600 while a 5090 lists $4,300-5,000. That price gap changes the answer, and the 70B question decides the rest.
Bottom Line
- Dual used RTX 3090s: $2,000-2,600 total (as of August 2026), 48GB VRAM, runs a 70B at Q4 at 16-22 tok/s. About 700W of GPU under load.
- One RTX 5090: $4,300-5,000 (as of August 2026), 32GB VRAM, roughly 3x the memory bandwidth per card. Cannot hold a dense 70B at Q4.
- The 2026 flip: in 2025 these cost about the same. The memory shortage roughly doubled both, but the 5090 doubled from a higher base — the dual build is now half the price.
- Decision rule: if you will actually run 70B-class dense models, buy the two 3090s. If your ceiling is 32B, buy neither at these prices before reading the counter-case below.
The comparison, at prices that are actually current
Most versions of this comparison were written when a used 3090 cost $700 and a 5090 cost $2,000. Neither number survived 2026. Live trackers in August 2026 put used 3090s at $1,000-1,300 (eBay fair range $1,202-1,296) and the RTX 5090 at $4,300-5,000+, with its $1,999 MSRP essentially fictional after three NVIDIA price hikes this year.
| Dual RTX 3090 (used) | RTX 5090 (new) | |
|---|---|---|
| Price, Aug 2026 | $2,000-2,600 | $4,300-5,000 |
| VRAM | 48GB (2x24GB) | 32GB |
| Memory bandwidth | ~936 GB/s per card | ~1,792 GB/s |
| Dense 70B at Q4 (~40GB weights) | Fits, with KV headroom | Does not fit |
| 70B Q4 speed | 16-22 tok/s | n/a on one card |
| GPU power under load | ~700W combined | up to ~575W |
| Warranty | None | Yes |
| Slots, PSU, complexity | 2 cards, 1200W-class PSU, spacing | 1 card, 1000W+ PSU |
The strange result: the price spike made the used option relatively better. Both roughly doubled, but doubling $2,200 hurts more than doubling $1,400. Anyone who dismissed the dual-3090 route in 2025 as “only slightly cheaper” should rerun the math.
What 48GB buys that 32GB cannot
A dense 70B at Q4_K_M is roughly 40GB of weights before KV cache. That number is the entire decision:
- 32GB (5090): a 70B fits only at Q2-class quants, where quality visibly degrades. Your comfortable ceiling is ~32B dense at Q4-Q5, or MoE models whose active weights stream well. Can 24GB run a 70B? walks the same wall at the tier below.
- 48GB (2x3090): 70B Q4 fits with room for real context. Community 2026 numbers land at 16-22 tok/s on Llama 70B Q4_K_M under llama.cpp — reading pace, fine for chat and agents, slow for bulk generation.
If your model list stops at 27-32B — and for agentic work in 2026 it often honestly does — the second 3090 buys you nothing. What local LLM fits my machine will tell you which side of that line you are on.
The two taxes on the dual build
Power and heat. ~700W of GPU under sustained load, plus platform. You need a quality 1200W-class power supply — the MSI MAG A1200PLS 1200W is an ATX 3.1 unit with native 12V-2x6 that covers this build. Buy a reputable unit, not the cheapest listing. You also need a case with real airflow, and tolerance for noise. At the 2026 US average of ~$0.18/kWh, heavy daily use adds tens of dollars a month; the electricity break-even math is here. Undervolting both cards costs almost no tokens per second on memory-bound inference and cuts the heat substantially.
Setup, minus a myth. You do not need NVLink, and you do not need a Threadripper. llama.cpp and Ollama split layers across GPUs over plain PCIe, and on consumer PCIe the layer-split (pipeline) mode carries less overhead than true tensor parallelism — NVLink mainly pays off for training and tensor-parallel serving stacks. Any board that gives two double-wide slots x8/x8 is enough. What you do inherit: two used cards means two chances of a bad card, so run the full used-3090 verification checklist on each, especially memtest_vulkan.
When the 5090 wins anyway
- Your ceiling is 32B. Then the 5090 is simply the faster, simpler, warrantied machine — the 48GB argument never activates.
- You generate in bulk. ~1,792 GB/s of bandwidth makes single-card generation on models that fit dramatically faster than anything the 3090s do.
- The box lives in your office. One modern card is quieter and cooler than two 350W space heaters from 2020.
But note what you pay for that at August 2026 prices: about $2,300 more than the dual build, for less total VRAM. If that stings, the 5090 vs 4090 vs used 3090 single-card comparison covers the middle paths.
The cards, if you have decided:
- EVGA RTX 3090 24GB — the 48GB route, two of these. Verify used cards individually before the return window closes.
- GIGABYTE RTX 5090 WINDFORCE 32GB — the single-card route, if 32GB covers your models.
- MSI MAG A1200PLS 1200W — the PSU for the dual-3090 route.
- Corsair RM1000e 1000W — enough for the single 5090 route.
Amazon affiliate links — we earn a small commission at no cost to you.
Sources
- ResalePrices — used RTX 3090 eBay price tracking, August 2026
- Tom’s Hardware — RTX 5090 street pricing, August 2026
- Best GPU for LLM — how to run two RTX 3090s for LLM inference in 2026
- Compute Market — multi-GPU local LLM setup guide 2026
- InsiderLLM — dual-GPU local AI setups 2026
See Also
- The Cheapest Way to Run a 70B Locally in 2026 — this build against every other 70B-capable box
- How to Buy a Used RTX 3090 Without Getting Burned — run this checklist twice
- RTX 5090 vs 4090 vs Used 3090 — the single-card version of this decision
- Best Local LLM for 64GB VRAM — the tier above, dual 5090s and beyond
- Local LLM Electricity Cost Break-Even — what 700W actually costs per month
- What PSU Do You Need for a Local AI Rig? — the full PSU sizing guide behind the 1200W call
Need OpenClaw fixed live?
Remote rescue sessions for gateway, auth, tunnel, VPS, and model access problems.
See Rescue Session