← All guides

Used EPYC Servers for CPU-Only MoE Inference in 2026

A used 8-channel EPYC server with 512GB of registered DDR4 will run a 300-460GB mixture-of-experts model that no consumer GPU can touch. Reported speeds sit at 4-6 tokens per second. The pitch you will read elsewhere is that used server RAM escaped the 2026 memory shortage. It did not escape; it escaped less. DDR4 ECC currently tracks at a median of $8.66 per gigabyte, so the 512GB that makes this build interesting costs about $4,400 on its own.

Bottom Line

  • It works, and it is slow. Reported CPU-only speeds on 8-channel EPYC: 4.2 tok/s on DeepSeek-R1 Q5_K_S (461.81GB), ~5.2 tok/s on GLM-5 Q3_K_XL (309GB). Usable for batch work, painful for chat.
  • Memory channels beat cores, and it is not close. Bandwidth is the constraint, not core count. Buy DIMMs before you buy cores.
  • CORRECTED 2026-08-25. This page originally read two cross-machine reports — a dual 96-core 7K62 at 2.9 tok/s against a single 48-core 7K62 at 4.2 tok/s — as proof that dual socket is slower. Those are two different machines, so the comparison cannot isolate the socket count. A same-machine A/B shows the second socket helping: 1.83x on a dense 70B, and 1.02x on DeepSeek R1. One socket is still the right buy for MoE work, for a better reason. See dual-socket vs single-socket EPYC for LLM inference.
  • The cheap-RAM premise is half true. DDR4 ECC RDIMM tracked at a $8.66/GB median against $36.20/GB for DDR5 when checked on 24 August 2026. That is a 4.2x relative advantage, but 512GB still costs roughly $2,300 at the cheapest listing and about $4,400 at the median.
  • The CPU is the cheap part. Used EPYC 7402 listings around $89; a used EPYC 7513 32-core around $379. The RAM costs five to fifty times the processor.
  • This only makes sense for models that fit nowhere else. If it fits in 24GB, a used RTX 3090 at $1,000-1,300 is cheaper and roughly six times faster.
  • Prompt processing is the hidden cost. Generation at 5 tok/s is tolerable. CPU prompt processing on a long context is not, and no build guide mentions it.
  • Plan for noise and a power bill. Server chassis fans and 200W-class CPUs are not home-office equipment.

Why Anyone Considers This

The models that matter most in 2026 stopped fitting on consumer hardware. As we set out in open weights are not local any more, frontier open-weight releases are now mixture-of-experts models in the 300GB to 700GB range at usable quantisation. No consumer GPU holds that. A 96GB workstation card does not hold that.

But MoE models have a property that rescues CPU inference: only a fraction of the parameters activate per token. A 671B model with roughly 37B active parameters does 37B-worth of arithmetic per token while needing all 671B resident in memory. That is a memory-capacity problem, not a compute problem, and system RAM is the cheapest capacity you can buy per gigabyte.

Hence the pitch: a used server with eight memory channels and 512GB of registered DDR4. We cover the small end of this idea in running a local LLM on 128GB of RAM with no GPU. This page is the large end, with the honest arithmetic.

The Measured Speeds

These figures come from user reports in the llama.cpp CPU-inference discussion. They are self-reported by individual owners, not a controlled benchmark, so read them as an order of magnitude rather than a spec sheet.

SystemMemoryModelSpeed
EPYC 7K62, 48 cores8×64GB (8-ch)DeepSeek-R1 Q5_K_S, 461.81GB4.2 tok/s
Dual EPYC 7K62, 96 cores16×64GBDeepSeek-R1 Q5_K_S, 461.81GB2.9 tok/s
Threadripper Pro 3955WX, 16 cores8×64GBDeepSeek-R1 Q5_K_S, 461.81GB2.8 tok/s
EPYC 7552, 48 cores512GB DDR4-2666 (8-ch)GLM-5 Q3_K_XL, 309GB~5.2 tok/s (24 threads)
EPYC 9654 (DDR5-4800)12-ch DDR5671B Q86.2 tok/s

Three things in that table deserve attention.

Row two carries a correction, added 2026-08-25. The dual-socket machine has twice the cores and twice the DIMMs, and it reports 31% slower than the single-socket machine on the identical workload. We originally concluded from that pair that dual socket is slower. That conclusion does not survive a controlled test. These are two different machines owned by two different people, configured differently, so the comparison cannot isolate the socket count. When a single machine is measured with one socket and then with both, the second socket is faster every time — although on DeepSeek R1 it is faster by only 2.2%. The full numbers, the mechanism, and the NUMA placement fix that recovers about 80% on a badly configured dual-socket box are in dual-socket vs single-socket EPYC for LLM inference. Buy one socket for MoE work — but buy it because the second socket returns almost nothing on MoE, not because it makes the machine slower.

Row three shows cores are not the lever either. A 16-core Threadripper Pro with the same 8 channels lands at 2.8 tok/s against 4.2 for a 48-core EPYC. There is some core scaling, but the EPYC 7552 report used only 24 of its 48 threads and still hit 5.2 tok/s. Past a point, adding threads adds contention, not speed.

Row five is the ceiling. Twelve channels of DDR5-4800 buys 6.2 tok/s. Even the modern, expensive version of this machine is not fast. Do not expect a newer platform to change the category.

The Real Cost of the RAM

This is where the topic usually gets oversold, so here are the numbers as checked on 24 August 2026.

A live server-memory tracker showed DDR4 ECC RDIMM at a $8.66/GB median and DDR5 at a $36.20/GB median, with the cheapest DDR4 listing at $4.53/GB for a 16GB DDR4-2133 RDIMM. Separately, refurbished-server-parts coverage of the 2026 market reports that refurbished DDR4 has itself risen 30-50%, while a new 64GB DDR5 RDIMM went from about $255 to over $900 in a year.

Applied to the capacity that makes this build worth doing:

CapacityAt cheapest listing ($4.53/GB)At median ($8.66/GB)Same capacity in DDR5 ($36.20/GB)
256GB~$1,160~$2,220~$9,270
512GB~$2,320~$4,430~$18,530
1TB~$4,640~$8,870~$37,070

Say the honest version out loud: the DDR5 column is why people call DDR4 cheap. Against DDR5 it looks like a bargain. Against a $1,150 used RTX 3090 it does not. Buying 512GB at the median price costs roughly what four used 3090s cost, and four 3090s give you 96GB of memory running at more than ten times the bandwidth.

Two practical notes on sourcing. The prices above come from a scraped tracker and a parts-reseller write-up, which is the weakest source class we use, so treat the ranges as indicative and price your actual DIMM configuration before committing. And the cheapest listing is a 16GB module: filling 8 channels to 512GB at $4.53/GB means 32 sticks, which most single-socket boards cannot hold. In practice you buy 64GB modules and pay closer to the median.

What the Rest of the Machine Costs

The processor is genuinely cheap, which is the part of the pitch that survives scrutiny. Recent listings show a used EPYC 7402 24-core around $89 and a used EPYC 7513 32-core around $379. Supermicro H12SSL-I boards for the 7002/7003 generation list around $1,178 new in Canadian dollars, with used and bundled options below that.

So a rough single-socket build, with every figure treated as indicative:

  • CPU: $89-379 used
  • Motherboard: several hundred used, roughly $900 US equivalent new
  • 512GB registered DDR4: $2,300-4,400
  • Chassis, PSU, cooling, storage: several hundred

The RAM is 60-80% of the bill. Every optimisation that matters is a memory-buying decision, not a compute decision.

When To Build This, and When Not To

Build it when all of these are true:

  1. The model you need is genuinely 200GB or larger at a quantisation you accept.
  2. It is a mixture-of-experts model with a small active-parameter fraction. A dense model of that size will be several times slower again.
  3. Your workload is batch or asynchronous. Overnight document processing, agent runs you do not watch, evaluation sweeps.
  4. You can put a loud machine somewhere that is not your office.

Do not build it when:

  • The model fits in 24GB or 48GB. Buy a card. It is cheaper and faster. Our 48GB setup guide covers that tier.
  • You want interactive chat. At 5 tok/s a 500-token answer takes over 90 seconds, and that ignores prompt processing.
  • Your context is long. CPU prompt processing scales badly and is the part these reports mostly do not measure. Assume it is worse than you hope, and test it before you buy.
  • You were sold on “cheap RAM.” Re-read the table above. It is cheaper than DDR5 and it is not cheap.
⚡ THE COMPARISON THAT MATTERS

Before you spend $3,000-5,000 on an EPYC build, check whether your model fits in 24GB. A used RTX 3090 costs $1,000-1,300 as of August 2026 and runs a 27B model at roughly 30 tok/s — about six times the speed of the server, for a quarter to a third of the price. The server wins only on capacity. We do not stock affiliate links for used EPYC parts, and we are not going to invent one; source those from eBay or a server-parts reseller directly.

Amazon affiliate link — we earn a small commission at no cost to you.

Making It As Fast As It Can Be

If you build it, these are the levers that actually move the number:

  1. One socket, all channels filled. Eight DIMMs on an 8-channel board. A half-populated board halves your bandwidth and therefore halves your speed.
  2. Fastest DIMM speed the platform supports. DDR4-3200 over DDR4-2666 where the CPU allows it. This is straight bandwidth.
  3. Fewer threads than you expect. The EPYC 7552 report used 24 of 48 threads. Sweep the thread count; the best value is often around half the physical cores.
  4. Use the MoE offload flags. Even a small GPU can hold the attention layers and the KV cache while system RAM holds the experts. Our guide to llama.cpp MoE offload flags covers the exact arguments, and running a 160GB MoE on 8GB of VRAM shows how far that goes.
  5. Quantise harder than feels comfortable. Q3 versus Q5 is the difference between fitting in 256GB and needing 512GB, and at these prices that is thousands of dollars.

See Also

Sources

Used RTX 3090 pricing is our own August 2026 price reference, compiled from eBay-derived trackers. The 30 tok/s comparison figure is the measured 27B result from the RTX 3090 power-limit sweep, not a like-for-like test against the server models above; it is there to show the order-of-magnitude difference, not to benchmark the same workload.

Need OpenClaw fixed live?

Remote rescue sessions for gateway, auth, tunnel, VPS, and model access problems.

See Rescue Session

Read next

Best Models for the Biggest Mac Studio: 96GB New, 256GB Used
Apple pulled the 512GB M3 Ultra in March 2026 and the 256GB in May — the biggest Mac Studio you can order new is 96GB. Best models for each tier: gpt-oss 120B (23-60 tok/s), Qwen3-VL 235B Q4 (~30 tok/s), GLM-4.7 358B Q3 (~15 tok/s), Llama 4 Maverick, and why DeepSeek V4 Flash finally runs local.
Best Local LLM for Mac Studio M3 Ultra (2026): 96GB New, 512GB Used
The best local LLM for the Mac Studio M3 Ultra at ~800 GB/s. Apple now sells only the 96GB configuration — the 512GB and 256GB options were pulled in 2026. Run 70B at Q8 and 100B+ MoE locally on what you can actually buy.
Best Local LLM for Mac Studio M5 Ultra (2026): 256GB
Apple announced the M5 Ultra Mac Studio on August 25, 2026 with 1.2TB/s bandwidth and a 256GB option — the first new 256GB Mac since Apple pulled the tier in May. What fits, projected tok/s, the $4,000 memory tax, and why you should wait for real benchmarks.
DGX Spark vs Mac Studio M3 Ultra for Local LLMs
Compare NVIDIA DGX Spark and Mac Studio M3 Ultra for local LLMs: 273 vs 819 GB/s bandwidth, prefill vs decode speed, 512GB memory, and 2026 pricing.