← All guides

Unified Memory vs PC RAM Explained (2026): Bandwidth

64GB is 64GB when you ask whether a model fits. It is not the same number when you ask how fast the model talks. To write each token, the machine reads every active weight from memory. A normal desktop reads its RAM at about 90 GB/s. A Mac mini M5 Pro reads its unified memory at 307 GB/s. That gap, not the capacity, is why the Mac feels usable and the PC feels stuck. This page gives the bandwidth of each memory type, the speed limit it sets, and a method to work out your own number.

Bottom Line

  • Capacity decides if the model loads. Bandwidth decides how fast it talks. Both a 64GB Mac and a 64GB PC load the same model. The Mac reads it three to seven times faster.
  • A desktop with two DDR5 sticks moves about 90 GB/s. DDR5-5600 x 8 bytes x 2 channels = 89.6 GB/s. A Mac mini M5 Pro moves 307 GB/s. A Mac Studio M5 Max moves 614 GB/s.
  • Your speed limit is bandwidth divided by the bytes read per token. A dense 27B model at Q4_K_M is 16.82GB. On dual-channel DDR5 that is a ceiling of about 5 tokens per second. On an M5 Pro it is about 18.
  • Mixture-of-experts models change the answer. They read only the active experts for each token. That is why 64GB of PC RAM runs a model like Qwen3.6-35B-A3B at usable speed.
  • A GPU is faster than any Mac, but only inside its VRAM. If the model spills out, the card reaches your RAM over PCIe at about 32 GB/s. That is slower than the CPU reading the same RAM.

The Decoder Table: Memory Type to Bandwidth to tok/s

The last column is a ceiling, not a benchmark. It is bandwidth divided by the size of one dense model file: Qwen3.6-27B at Q4_K_M, 16.82GB on Hugging Face (read 2026-09-24). Real runs land below it. See the receipts further down.

MemoryWhere you find itBandwidth (GB/s)How we got the numberCeiling, dense 27B Q4_K_M (tok/s)
DDR5-5600, 2 channelsMost desktops (2 or 4 sticks)89.6Computed: 5,600 MT/s x 8 bytes x 25.3
DDR5-6000, 2 channelsTuned gaming desktops96.0Computed: 6,000 x 8 x 25.7
DDR5-6400, 4 channelsThreadripper on TRX50204.8Computed: 6,400 x 8 x 412.2
DDR5-6400, 8 channelsThreadripper PRO 9000 WX409.6Computed: 6,400 x 8 x 824.4
Unified, M6 16GBMac mini M6 base153Apple Mac mini specs9.1 (tight fit)
Unified, M6 24GB / 32GBMac mini M6170Apple Mac mini specs10.1
Unified, M5 ProMac mini M5 Pro307Apple Mac mini specs18.3
Unified, M5 Max 32-core GPUMac Studio460Apple Mac Studio specs27.3
Unified, M5 Max 40-core GPUMac Studio614Apple Mac Studio specs36.5
Unified, M5 UltraMac Studio1,200Apple lists 1.2 TB/s71.3
LPDDR5X-8000, 256-bitStrix Halo (Ryzen AI Max+ 395)256AMD: 256 GB/s; computed 8,000 x 32 bytes15.2
LPDDR5X, unifiedNVIDIA DGX Spark273NVIDIA DGX Spark page16.2
GDDR7, 128-bitRTX 5060 Ti 16GB44828 Gbps x 128 bits / 8Does not fit (16.82GB file, 16GB card)
GDDR6X, 384-bitRTX 3090 24GB936NVIDIA GA102 whitepaper55.6
GDDR7, 512-bitRTX 5090 32GB1,792NVIDIA RTX Blackwell whitepaper106.5
PCIe 4.0 x16 linkGPU reading system RAM~31.5Computed: 16 GT/s x 16 lanes x 128/130 / 8The spill path, see below
PCIe 5.0 x16 linkGPU reading system RAM~63Computed: 32 GT/s x 16 lanes x 128/130 / 8The spill path, see below

Read the table in two passes. First, find a row with enough capacity to hold your model. Then read its bandwidth. The first pass tells you if the model runs. The second tells you if you will want to use it.

Macs give the GPU part of unified memory, not all of it. The community-reported default is about 75%, so a 64GB Mac has about 48GB for models. Treat that as a community figure, not an Apple spec. On Strix Halo, AMD says up to 96GB of a 128GB machine can be set as graphics memory.

Mechanism 1: Capacity Decides Whether the Model Loads

A model file must sit in memory before it runs. Add the KV cache for your context on top. If the total is larger than the memory the processor can use, the model does not load, or it pages to disk and crawls.

On this question a 64GB PC and a 64GB Mac are close. The PC gives the CPU nearly all 64GB. The Mac gives the GPU about 48GB by default. For capacity alone, the PC is not worse.

Mechanism 2: Bandwidth Decides Tokens per Second

Token generation (decode) is one pass through the model for each token. Each pass reads every active weight from memory. The processor waits on memory, not on math. So the ceiling is:

tok/s ceiling = memory bandwidth (GB/s) / bytes read per token (GB)

For a dense model, bytes read per token is close to the file size. Dual-channel DDR5 at 89.6 GB/s divided by 16.82GB is 5.3 tokens per second. The M5 Pro at 307 GB/s gives 18.3. Same capacity class, three times the speed.

Prompt processing (prefill) is different. It is compute-bound, so GPU cores matter there. That is why Apple’s larger GPU options speed up long prompts but do not raise token speed when the bandwidth figure is the same.

Why a PC’s GPU Cannot Use System RAM at VRAM Speed

A discrete GPU reads its own VRAM at 448 to 1,792 GB/s. It reaches system RAM through the PCIe slot. PCIe 4.0 x16 carries about 31.5 GB/s in each direction. PCIe 5.0 x16 carries about 63 GB/s.

The RTX 5060 Ti uses a PCIe 5.0 x8 link. That is about 31.5 GB/s, the same as PCIe 4.0 x16. On a PCIe 4.0 board, the x8 link falls to about 15.8 GB/s (computed).

So when a model spills out of VRAM, the spilled part runs at PCIe speed. That is slower than dual-channel DDR5 read by the CPU. This is why llama.cpp computes offloaded layers on the CPU, next to the RAM, instead of streaming them to the GPU. It is also why speed falls off a cliff at the exact point VRAM fills. If you are sizing a card for this, see the cards this matters for.

A Mac, a Strix Halo box and a DGX Spark have no such seam. The CPU and GPU share one memory pool at one bandwidth. That is the whole meaning of “unified” for local AI.

The MoE Exception: Why 64GB of PC RAM Still Works

A Mixture-of-Experts model stores many experts but uses a few per token. The bytes read per token follow the active parameters, not the total.

ModelTotal / active paramsFileBytes read per token (estimate)Ceiling on 89.6 GB/s DDR5
Qwen3.6-27B (dense)27B / 27B16.82GB Q4_K_M~16.8GB~5 tok/s
Qwen3.6-35B-A3B (MoE)35B / 3B22.13GB UD-Q4_K_M~1.9GB~47 tok/s
gpt-oss 120B (MoE)117B / 5.1B65.25GBnot estimated heredoes not fit 64GB with context

The estimate for Qwen3.6-35B-A3B is the file size scaled by 3B / 35B. It ignores that some tensors use a higher type, so the real figure is higher and the real speed is lower. CPU compute also caps small active sets before bandwidth does. Treat 47 as a ceiling you will not reach on a CPU.

The practical move on a PC with a GPU is to split the model. Keep attention on the GPU and put the expert weights in system RAM. The llama.cpp MoE offload flags page shows the flags and the sweep. The 64GB RAM model list shows which MoE models fit.

Receipts: Measured Runs vs the Ceiling

Apple Silicon, llama.cpp discussion #4167 (community-reported). The table runs llama-bench on Llama 2 7B at Q4_0, a 3.56 GiB file (3.82GB). Text generation at batch size 1:

ChipBandwidthCeiling (derived)Measured TGShare of ceiling
M4, 10-core GPU120 GB/s31.4 tok/s24.11 tok/s77%
M4 Pro, 20-core GPU273 GB/s71.4 tok/s50.74 tok/s71%
M5 Pro, 20-core GPU307 GB/s80.3 tok/s66.33 tok/s83%
M5 Max, 40-core GPU614 GB/s160.6 tok/s119.92 tok/s75%
M2 Ultra, 76-core GPU800 GB/s209.3 tok/s94.27 tok/s45%

Speed follows bandwidth closely up to the Max chips. The Ultra reaches less than half of its ceiling on a small model. More bandwidth is not free speed at the top end.

A DDR5 desktop that ran at DDR4 speed, llama.cpp issue #4716 (community-reported). A Ryzen 9 7950X with 128GB of DDR5 at 3,600 MT/s ran Mixtral Q8_0 at 3.37 tok/s. A Ryzen 9 5900X with 128GB of DDR4 at the same 3,600 MT/s ran it at 3.58 tok/s. The newer CPU did not help, because both machines had the same memory speed. Our derived bandwidth is 3,600 x 8 x 2 = 57.6 GB/s. Mixtral reads about 13.7GB per token at Q8_0 (12.9B active params x 8.5 bits / 8). That is 4.2 tok/s, and the 7950X reached 80% of it.

That report holds the one fact most “64GB PC” advice misses. The user filled all four slots to reach 128GB, and the DDR5 ran at 3,600 MT/s instead of a 5,600 or 6,000 rating. Rated speed is the speed of the kit, not the speed your board runs with four sticks. Check the speed your BIOS reports.

Method: Find Your Own Number

The numbers above do not transfer between machines. Measure yours in four steps.

  1. Get your real memory speed. On Windows, Task Manager > Performance > Memory shows “Speed”. On Linux, run sudo dmidecode -t memory | grep -i "configured memory speed". On a Mac, use the bandwidth on Apple’s specs page for your chip.
  2. Compute bandwidth. For a desktop: MT/s x 8 x channels / 1,000 = GB/s. Two or four sticks on a mainstream desktop is still two channels.
  3. Compute bytes per token. Dense model: use the file size. MoE model: file size x active params / total params, then round up. The GGUF quant names decoder shows how to read the real file size.
  4. Measure and compare. Run llama-bench -m model.gguf -p 0 -n 128 and read the tg128 line. Divide it by your ceiling. Expect somewhere between 45% and 83%, the range in the Apple table above. The llama.cpp flags guide covers the flags that change it.

If your share is far below that range, suspect a VRAM spill, a model paged to disk, or memory running below its rated speed.

What the Table Means for a Purchase

  • You have a desktop and 64GB of DDR5: run MoE models. Dense 27B-class models will run at about 4 to 5 tok/s. Add a GPU for attention and use the offload flags.
  • You want dense 27B-class models at reading speed on a small box: the Mac mini M5 Pro at 307 GB/s is the first current Mac mini above 300 GB/s. The Mac mini config picker prices each option.
  • You want 128GB of fast memory under the price of a Mac Studio 128GB: Strix Halo gives 256 GB/s with up to 96GB as graphics memory. A Ryzen AI Max+ 395 128GB mini-PC is the box to check. The Strix Halo box picker compares the vendors.
  • The model fits in 24GB or 32GB: a GPU wins on speed. A 3090 has 936 GB/s. No Mac mini comes close.

One 2026 note on cost. The DRAM shortage moved the price of PC capacity. Newegg listed a 64GB DDR5-6000 kit at $869.99 on 2026-07-31, roughly four times the pre-shortage price. We found no lower September 2026 quote; check current listings. Capacity used to be the cheap half of a PC. It no longer is, and it still does not buy bandwidth.

Common Mistakes

  1. Comparing capacity only. 64GB and 64GB tells you both machines load the model. It says nothing about speed.
  2. Counting sticks as channels. Four sticks on a mainstream desktop are still two channels, and they can run slower than two. The #4716 machine ran at 3,600 MT/s.
  3. Assuming the GPU can use system RAM as VRAM. It can, over PCIe, at about 32 GB/s. That is slower than the CPU on the same RAM.
  4. Using total parameters for an MoE model. Use the active parameters for speed. Use the total for capacity.
  5. Using Apple’s “4x faster” LLM claims for token speed. Those figures describe prompt processing. Token speed follows the bandwidth figure.

FAQ

If 64GB of unified RAM on a MacBook can run good LLMs, why can't 64GB of RAM on a PC do that?

It can load the same models. It runs them slower, because of memory bandwidth. To write each token, the machine reads every active weight from memory. Dual-channel DDR5-5600 on a desktop moves 89.6 GB/s (5,600 MT/s x 8 bytes x 2 channels). Apple lists 307 GB/s for the M5 Pro and 614 GB/s for the 40-core M5 Max. For a dense 27B model at Q4_K_M (16.82GB), the ceiling is about 5 tokens per second on the PC. On the M5 Pro it is about 18. Mixture-of-experts models read only their active experts per token, so 64GB of PC RAM runs them at usable speed.

Mac mini vs GPU for LLM: which is faster?

A GPU is faster for any model that fits in its VRAM. NVIDIA lists 1,792 GB/s for the RTX 5090 and 936 GB/s for the RTX 3090. Apple lists 170 GB/s for the 24GB and 32GB Mac mini M6 and 307 GB/s for the M5 Pro. The Mac mini wins on capacity. The M5 Pro Mac mini takes up to 64GB, which no 24GB or 32GB card can hold. When a model spills out of VRAM, the GPU reaches system RAM over PCIe, about 32 GB/s on PCIe 4.0 x16, and speed collapses.

Mac mini vs Mac Studio for local LLM: what is the difference?

Bandwidth and the memory ceiling. Apple lists 170 GB/s for the 24GB and 32GB Mac mini M6 and 153 GB/s for the 16GB. The Mac mini M5 Pro reads memory at 307 GB/s and tops out at 64GB. The Mac Studio M5 Max reads memory at 460 GB/s (32-core GPU) or 614 GB/s (40-core GPU) and goes to 128GB. The M5 Ultra reads it at 1.2 TB/s. The same model generates tokens about twice as fast on a 614 GB/s Mac Studio as on a 307 GB/s Mac mini.

AMD Ryzen AI Halo vs Mac Studio: which is faster for LLMs?

The Mac Studio, on token generation. The Ryzen AI Max+ 395 (Strix Halo, sold in the Ryzen AI Halo box) uses LPDDR5X-8000 on a 256-bit bus, which is 256 GB/s. Apple lists 460 GB/s or 614 GB/s for the M5 Max Mac Studio and 1.2 TB/s for the M5 Ultra. Strix Halo wins on memory per dollar: 128GB, of which AMD says up to 96GB can be set as graphics memory. Both hold models that no single consumer GPU holds.

See Also

Sources

Need OpenClaw fixed live?

Remote rescue sessions for gateway, auth, tunnel, VPS, and model access problems.

See Rescue Session

Read next

RAM Prices for Local AI, September 2026: 64GB DDR5 Now Costs More Than a PS5
A 64GB DDR5 kit sits near $1,099-1,272 in September 2026, up roughly 473-485% from about $222. Current prices per capacity, what it does to a local AI build budget, and the three buying paths that still work.
Best Local LLM for Mac mini (2026): Why 24GB Is the Floor
Best local LLM for the Mac mini in 2026, by memory tier. The M6 base has 16GB (about 12GB usable) and runs Qwen 3.5 9B Q8_0. The 32GB M6 and 64GB M5 Pro restore the ceilings Apple deleted in May 2026. The outgoing M4 is a value pick only under about $705.
The Soldered Memory Trap: Why 'Buy Less Now, Upgrade Later' Fails on Unified-Memory AI Boxes
Macs, Strix Halo mini-PCs and the DGX Spark all solder their memory. You buy your RAM ceiling once, permanently. Worse, in 2026 vendors deleted configs mid-generation — Apple removed the 64GB Mac mini M4 Pro and the 256GB/512GB Mac Studio. The config you planned to upgrade to may not exist when you go back.
Which Mac mini Should You Buy for Local LLMs? (2026)
The exact Mac mini configuration to order for local LLMs, as of September 2026. Nine M6 and M5 Pro configs from $899 to $2,899, priced per GB. Buy the 32GB M6 at $1,299. The $899 16GB base runs memory at 153 GB/s, not 170, and no Mac mini runs gpt-oss 120B.