Should I Wait for the RTX 60 Series? (August 2026 Answer: No)
The honest answer is no, and the reason is not that the RTX 60 series will be bad. It is that the wait is far longer than most people think, and the thing you would normally do while waiting — pick up a discounted last-gen card — does not work in 2026. Leaks put Rubin-based GeForce at second-half 2027, possibly 2028. The RTX 50 Super refresh that would have filled the gap was shelved indefinitely. And last-gen cards are getting more expensive, not less.
Bottom Line (August 2026)
- The wait is 12–18 months minimum. RTX 60 (Rubin, GR20x) is rumoured for 2H 2027, with credible reporting that it slips to 2028. Unconfirmed by NVIDIA.
- The bridge product is gone. The RTX 50 SUPER refresh was delayed indefinitely. The 24GB RTX 5080 Super reached completed-design stage and was shelved.
- “Wait and buy last-gen cheap” is broken in 2026. Several last-gen cards trade above their own launch MSRP. The RTX 4070 Ti Super launched at $799 and sold at ~$1,100–1,200 in August 2026.
- Expected gain is ~30% on the flagship, per leaks. That is a normal generation, not a reason to skip 18 months.
- Buy the VRAM you need now. For local LLMs, capacity decides what runs; generation only decides how fast.
Why “Just Wait” Fails This Time
Waiting for the next generation is usually a reasonable strategy. It rests on two assumptions that both held for a decade:
- The next generation arrives in roughly 24 months.
- While you wait, the current generation gets cheaper.
In August 2026, both assumptions are false.
The timeline is longer than a normal cycle
Reporting from multiple outlets, tracing primarily to leaker kopite7kimi, puts RTX 60 in the second half of 2027, using the Rubin architecture with a GR20x GPU family. Several of those same reports flag a possible slip into 2028.
The stated reasons matter more than the date, because they tell you whether it will slip again:
- GDDR7 supply is tight. Memory is the constraint, not the GPU die.
- NVIDIA is prioritising datacenter. Limited manufacturing and HBM allocation goes to AI accelerators with far better margins than GeForce.
- The RTX 50 SUPER delay cascades. Removing the mid-cycle refresh shifted the whole roadmap.
None of those resolve on a schedule you can see. A date driven by a memory shortage is a date that slips when the shortage does not.
The refresh you were actually waiting for is cancelled
Most people saying “I’ll wait” in 2026 are not waiting for RTX 60. They are waiting for the RTX 50 SUPER refresh — a 12-month wait rather than a 24-month one.
That product is off the table. NVIDIA informed partners it is delayed indefinitely, and the RTX 5080 Super with 24GB reached completed-design stage before being shelved. AIC allocations were cut 15–20%.
For local AI specifically, that shelved 24GB 5080 was the single most interesting unreleased card on the roadmap. A 24GB card at 5080 pricing would have reset the VRAM-per-dollar ladder. Its cancellation is the reason “wait for the Super” stopped being advice and became just delay.
Last-gen is getting more expensive
This is the part that breaks people’s mental model, and it is worth stating bluntly: in 2026, the “last generation gets cheaper” rule is false.
| Card | Launch MSRP | Street price, Aug 2026 |
|---|---|---|
| RTX 4070 Ti Super 16GB | $799 | ~$1,100–1,200 |
| RTX 3090 24GB (used) | $1,499 | $1,000–1,300 |
| RTX 5090 32GB | $1,999 | $4,300–5,000+ |
A card launched at $799 selling for $1,100–1,200 two years later is not a market where waiting pays. NVIDIA ran three GeForce price increases in 2026 (January, May, and late July/August). Memory is now reported at over 80% of the bill of materials on some high-end cards.
The line nobody else prints: waiting is not free and it is not neutral — in a rising market it is a position. Every month you wait, the card you eventually buy has cost you more, and the card you would have used in the meantime has appreciated in your hands. This is the first GPU cycle in memory where the person who bought early is ahead of the person who waited.
What to Buy Instead of Waiting
The rule for local AI is simple: VRAM capacity decides what you can run at all; generation decides how fast. A 2020 card with 24GB runs models a 2026 card with 12GB cannot load at any speed.
| Your situation | Buy this now |
|---|---|
| Running 7B–14B models, tight budget | Intel Arc B580 12GB (~$300–310) |
| Want the 16GB new-card floor | RTX 5060 Ti 16GB |
| Serious local inference, best $/GB | Used RTX 3090 24GB — still the value floor |
| Best $/GB on a new card | RX 7900 XTX 24GB (~$900–950, near MSRP) |
| Need 32GB and accept the price | RTX 5090 32GB |
The RX 7900 XTX is worth a second look in this market. It is one of the rare cards still trading at or below MSRP, which currently makes it the best VRAM-per-dollar among new consumer cards. The tradeoff is real: ROCm is a less mature stack than CUDA, and some tooling assumes NVIDIA.
When Waiting Is Correct
To be fair to the other side, there are three cases where waiting genuinely wins:
- Your current hardware already runs what you need. If a 16GB card serves your models fine, there is no reason to buy into a shortage. Waiting is free when you are not blocked.
- You need something that does not exist yet. If your requirement is 48GB on a single consumer card, no purchase available today satisfies it. Wait, or move to a workstation card.
- You can rent instead. Bursty workloads — occasional fine-tuning, a one-off batch job — are cheaper rented than owned at 2026 prices. See our fine-tuning vs inference guide for where that line sits.
Outside those three, buying the smallest card that clears your VRAM requirement today beats waiting 18 months for a card whose launch price will be set by the same memory shortage that is inflating everything now.
See Also
- The RTX 50 Super Is Cancelled. What Should You Buy Instead? — the full cancellation story
- Best GPU for Fine-Tuning vs Inference — when to buy versus rent
- RTX 5090 vs 4090 vs Used 3090 — the three-way buying comparison
- Cheapest 32GB VRAM GPU — if capacity is your binding constraint
- Best Local LLM by GPU (hub) — per-GPU model picks
- The Soldered Memory Trap — the same “buy now vs later” question on unified-memory boxes
Need OpenClaw fixed live?
Remote rescue sessions for gateway, auth, tunnel, VPS, and model access problems.
See Rescue Session