Unified Memory vs PC RAM Explained (2026): Bandwidth
64GB is 64GB when you ask whether a model fits. It is not the same number when you ask how fast the model talks. To write each token, the machine reads every active weight from memory. A normal desktop reads its RAM at about 90 GB/s. A Mac mini M5 Pro reads its unified memory at 307 GB/s. That gap, not the capacity, is why the Mac feels usable and the PC feels stuck. This page gives the bandwidth of each memory type, the speed limit it sets, and a method to work out your own number.
Bottom Line
- Capacity decides if the model loads. Bandwidth decides how fast it talks. Both a 64GB Mac and a 64GB PC load the same model. The Mac reads it three to seven times faster.
- A desktop with two DDR5 sticks moves about 90 GB/s. DDR5-5600 x 8 bytes x 2 channels = 89.6 GB/s. A Mac mini M5 Pro moves 307 GB/s. A Mac Studio M5 Max moves 614 GB/s.
- Your speed limit is bandwidth divided by the bytes read per token. A dense 27B model at Q4_K_M is 16.82GB. On dual-channel DDR5 that is a ceiling of about 5 tokens per second. On an M5 Pro it is about 18.
- Mixture-of-experts models change the answer. They read only the active experts for each token. That is why 64GB of PC RAM runs a model like Qwen3.6-35B-A3B at usable speed.
- A GPU is faster than any Mac, but only inside its VRAM. If the model spills out, the card reaches your RAM over PCIe at about 32 GB/s. That is slower than the CPU reading the same RAM.
The Decoder Table: Memory Type to Bandwidth to tok/s
The last column is a ceiling, not a benchmark. It is bandwidth divided by the size of one dense model file: Qwen3.6-27B at Q4_K_M, 16.82GB on Hugging Face (read 2026-09-24). Real runs land below it. See the receipts further down.
| Memory | Where you find it | Bandwidth (GB/s) | How we got the number | Ceiling, dense 27B Q4_K_M (tok/s) |
|---|---|---|---|---|
| DDR5-5600, 2 channels | Most desktops (2 or 4 sticks) | 89.6 | Computed: 5,600 MT/s x 8 bytes x 2 | 5.3 |
| DDR5-6000, 2 channels | Tuned gaming desktops | 96.0 | Computed: 6,000 x 8 x 2 | 5.7 |
| DDR5-6400, 4 channels | Threadripper on TRX50 | 204.8 | Computed: 6,400 x 8 x 4 | 12.2 |
| DDR5-6400, 8 channels | Threadripper PRO 9000 WX | 409.6 | Computed: 6,400 x 8 x 8 | 24.4 |
| Unified, M6 16GB | Mac mini M6 base | 153 | Apple Mac mini specs | 9.1 (tight fit) |
| Unified, M6 24GB / 32GB | Mac mini M6 | 170 | Apple Mac mini specs | 10.1 |
| Unified, M5 Pro | Mac mini M5 Pro | 307 | Apple Mac mini specs | 18.3 |
| Unified, M5 Max 32-core GPU | Mac Studio | 460 | Apple Mac Studio specs | 27.3 |
| Unified, M5 Max 40-core GPU | Mac Studio | 614 | Apple Mac Studio specs | 36.5 |
| Unified, M5 Ultra | Mac Studio | 1,200 | Apple lists 1.2 TB/s | 71.3 |
| LPDDR5X-8000, 256-bit | Strix Halo (Ryzen AI Max+ 395) | 256 | AMD: 256 GB/s; computed 8,000 x 32 bytes | 15.2 |
| LPDDR5X, unified | NVIDIA DGX Spark | 273 | NVIDIA DGX Spark page | 16.2 |
| GDDR7, 128-bit | RTX 5060 Ti 16GB | 448 | 28 Gbps x 128 bits / 8 | Does not fit (16.82GB file, 16GB card) |
| GDDR6X, 384-bit | RTX 3090 24GB | 936 | NVIDIA GA102 whitepaper | 55.6 |
| GDDR7, 512-bit | RTX 5090 32GB | 1,792 | NVIDIA RTX Blackwell whitepaper | 106.5 |
| PCIe 4.0 x16 link | GPU reading system RAM | ~31.5 | Computed: 16 GT/s x 16 lanes x 128/130 / 8 | The spill path, see below |
| PCIe 5.0 x16 link | GPU reading system RAM | ~63 | Computed: 32 GT/s x 16 lanes x 128/130 / 8 | The spill path, see below |
Read the table in two passes. First, find a row with enough capacity to hold your model. Then read its bandwidth. The first pass tells you if the model runs. The second tells you if you will want to use it.
Macs give the GPU part of unified memory, not all of it. The community-reported default is about 75%, so a 64GB Mac has about 48GB for models. Treat that as a community figure, not an Apple spec. On Strix Halo, AMD says up to 96GB of a 128GB machine can be set as graphics memory.
Mechanism 1: Capacity Decides Whether the Model Loads
A model file must sit in memory before it runs. Add the KV cache for your context on top. If the total is larger than the memory the processor can use, the model does not load, or it pages to disk and crawls.
On this question a 64GB PC and a 64GB Mac are close. The PC gives the CPU nearly all 64GB. The Mac gives the GPU about 48GB by default. For capacity alone, the PC is not worse.
Mechanism 2: Bandwidth Decides Tokens per Second
Token generation (decode) is one pass through the model for each token. Each pass reads every active weight from memory. The processor waits on memory, not on math. So the ceiling is:
tok/s ceiling = memory bandwidth (GB/s) / bytes read per token (GB)
For a dense model, bytes read per token is close to the file size. Dual-channel DDR5 at 89.6 GB/s divided by 16.82GB is 5.3 tokens per second. The M5 Pro at 307 GB/s gives 18.3. Same capacity class, three times the speed.
Prompt processing (prefill) is different. It is compute-bound, so GPU cores matter there. That is why Apple’s larger GPU options speed up long prompts but do not raise token speed when the bandwidth figure is the same.
Why a PC’s GPU Cannot Use System RAM at VRAM Speed
A discrete GPU reads its own VRAM at 448 to 1,792 GB/s. It reaches system RAM through the PCIe slot. PCIe 4.0 x16 carries about 31.5 GB/s in each direction. PCIe 5.0 x16 carries about 63 GB/s.
The RTX 5060 Ti uses a PCIe 5.0 x8 link. That is about 31.5 GB/s, the same as PCIe 4.0 x16. On a PCIe 4.0 board, the x8 link falls to about 15.8 GB/s (computed).
So when a model spills out of VRAM, the spilled part runs at PCIe speed. That is slower than dual-channel DDR5 read by the CPU. This is why llama.cpp computes offloaded layers on the CPU, next to the RAM, instead of streaming them to the GPU. It is also why speed falls off a cliff at the exact point VRAM fills. If you are sizing a card for this, see the cards this matters for.
A Mac, a Strix Halo box and a DGX Spark have no such seam. The CPU and GPU share one memory pool at one bandwidth. That is the whole meaning of “unified” for local AI.
The MoE Exception: Why 64GB of PC RAM Still Works
A Mixture-of-Experts model stores many experts but uses a few per token. The bytes read per token follow the active parameters, not the total.
| Model | Total / active params | File | Bytes read per token (estimate) | Ceiling on 89.6 GB/s DDR5 |
|---|---|---|---|---|
| Qwen3.6-27B (dense) | 27B / 27B | 16.82GB Q4_K_M | ~16.8GB | ~5 tok/s |
| Qwen3.6-35B-A3B (MoE) | 35B / 3B | 22.13GB UD-Q4_K_M | ~1.9GB | ~47 tok/s |
| gpt-oss 120B (MoE) | 117B / 5.1B | 65.25GB | not estimated here | does not fit 64GB with context |
The estimate for Qwen3.6-35B-A3B is the file size scaled by 3B / 35B. It ignores that some tensors use a higher type, so the real figure is higher and the real speed is lower. CPU compute also caps small active sets before bandwidth does. Treat 47 as a ceiling you will not reach on a CPU.
The practical move on a PC with a GPU is to split the model. Keep attention on the GPU and put the expert weights in system RAM. The llama.cpp MoE offload flags page shows the flags and the sweep. The 64GB RAM model list shows which MoE models fit.
Receipts: Measured Runs vs the Ceiling
Apple Silicon, llama.cpp discussion #4167 (community-reported). The table runs llama-bench on Llama 2 7B at Q4_0, a 3.56 GiB file (3.82GB). Text generation at batch size 1:
| Chip | Bandwidth | Ceiling (derived) | Measured TG | Share of ceiling |
|---|---|---|---|---|
| M4, 10-core GPU | 120 GB/s | 31.4 tok/s | 24.11 tok/s | 77% |
| M4 Pro, 20-core GPU | 273 GB/s | 71.4 tok/s | 50.74 tok/s | 71% |
| M5 Pro, 20-core GPU | 307 GB/s | 80.3 tok/s | 66.33 tok/s | 83% |
| M5 Max, 40-core GPU | 614 GB/s | 160.6 tok/s | 119.92 tok/s | 75% |
| M2 Ultra, 76-core GPU | 800 GB/s | 209.3 tok/s | 94.27 tok/s | 45% |
Speed follows bandwidth closely up to the Max chips. The Ultra reaches less than half of its ceiling on a small model. More bandwidth is not free speed at the top end.
A DDR5 desktop that ran at DDR4 speed, llama.cpp issue #4716 (community-reported). A Ryzen 9 7950X with 128GB of DDR5 at 3,600 MT/s ran Mixtral Q8_0 at 3.37 tok/s. A Ryzen 9 5900X with 128GB of DDR4 at the same 3,600 MT/s ran it at 3.58 tok/s. The newer CPU did not help, because both machines had the same memory speed. Our derived bandwidth is 3,600 x 8 x 2 = 57.6 GB/s. Mixtral reads about 13.7GB per token at Q8_0 (12.9B active params x 8.5 bits / 8). That is 4.2 tok/s, and the 7950X reached 80% of it.
That report holds the one fact most “64GB PC” advice misses. The user filled all four slots to reach 128GB, and the DDR5 ran at 3,600 MT/s instead of a 5,600 or 6,000 rating. Rated speed is the speed of the kit, not the speed your board runs with four sticks. Check the speed your BIOS reports.
Method: Find Your Own Number
The numbers above do not transfer between machines. Measure yours in four steps.
- Get your real memory speed. On Windows, Task Manager > Performance > Memory shows “Speed”. On Linux, run
sudo dmidecode -t memory | grep -i "configured memory speed". On a Mac, use the bandwidth on Apple’s specs page for your chip. - Compute bandwidth. For a desktop: MT/s x 8 x channels / 1,000 = GB/s. Two or four sticks on a mainstream desktop is still two channels.
- Compute bytes per token. Dense model: use the file size. MoE model: file size x active params / total params, then round up. The GGUF quant names decoder shows how to read the real file size.
- Measure and compare. Run
llama-bench -m model.gguf -p 0 -n 128and read the tg128 line. Divide it by your ceiling. Expect somewhere between 45% and 83%, the range in the Apple table above. The llama.cpp flags guide covers the flags that change it.
If your share is far below that range, suspect a VRAM spill, a model paged to disk, or memory running below its rated speed.
What the Table Means for a Purchase
- You have a desktop and 64GB of DDR5: run MoE models. Dense 27B-class models will run at about 4 to 5 tok/s. Add a GPU for attention and use the offload flags.
- You want dense 27B-class models at reading speed on a small box: the Mac mini M5 Pro at 307 GB/s is the first current Mac mini above 300 GB/s. The Mac mini config picker prices each option.
- You want 128GB of fast memory under the price of a Mac Studio 128GB: Strix Halo gives 256 GB/s with up to 96GB as graphics memory. A Ryzen AI Max+ 395 128GB mini-PC is the box to check. The Strix Halo box picker compares the vendors.
- The model fits in 24GB or 32GB: a GPU wins on speed. A 3090 has 936 GB/s. No Mac mini comes close.
One 2026 note on cost. The DRAM shortage moved the price of PC capacity. Newegg listed a 64GB DDR5-6000 kit at $869.99 on 2026-07-31, roughly four times the pre-shortage price. We found no lower September 2026 quote; check current listings. Capacity used to be the cheap half of a PC. It no longer is, and it still does not buy bandwidth.
Common Mistakes
- Comparing capacity only. 64GB and 64GB tells you both machines load the model. It says nothing about speed.
- Counting sticks as channels. Four sticks on a mainstream desktop are still two channels, and they can run slower than two. The #4716 machine ran at 3,600 MT/s.
- Assuming the GPU can use system RAM as VRAM. It can, over PCIe, at about 32 GB/s. That is slower than the CPU on the same RAM.
- Using total parameters for an MoE model. Use the active parameters for speed. Use the total for capacity.
- Using Apple’s “4x faster” LLM claims for token speed. Those figures describe prompt processing. Token speed follows the bandwidth figure.
FAQ
If 64GB of unified RAM on a MacBook can run good LLMs, why can't 64GB of RAM on a PC do that?
It can load the same models. It runs them slower, because of memory bandwidth. To write each token, the machine reads every active weight from memory. Dual-channel DDR5-5600 on a desktop moves 89.6 GB/s (5,600 MT/s x 8 bytes x 2 channels). Apple lists 307 GB/s for the M5 Pro and 614 GB/s for the 40-core M5 Max. For a dense 27B model at Q4_K_M (16.82GB), the ceiling is about 5 tokens per second on the PC. On the M5 Pro it is about 18. Mixture-of-experts models read only their active experts per token, so 64GB of PC RAM runs them at usable speed.
Mac mini vs GPU for LLM: which is faster?
A GPU is faster for any model that fits in its VRAM. NVIDIA lists 1,792 GB/s for the RTX 5090 and 936 GB/s for the RTX 3090. Apple lists 170 GB/s for the 24GB and 32GB Mac mini M6 and 307 GB/s for the M5 Pro. The Mac mini wins on capacity. The M5 Pro Mac mini takes up to 64GB, which no 24GB or 32GB card can hold. When a model spills out of VRAM, the GPU reaches system RAM over PCIe, about 32 GB/s on PCIe 4.0 x16, and speed collapses.
Mac mini vs Mac Studio for local LLM: what is the difference?
Bandwidth and the memory ceiling. Apple lists 170 GB/s for the 24GB and 32GB Mac mini M6 and 153 GB/s for the 16GB. The Mac mini M5 Pro reads memory at 307 GB/s and tops out at 64GB. The Mac Studio M5 Max reads memory at 460 GB/s (32-core GPU) or 614 GB/s (40-core GPU) and goes to 128GB. The M5 Ultra reads it at 1.2 TB/s. The same model generates tokens about twice as fast on a 614 GB/s Mac Studio as on a 307 GB/s Mac mini.
AMD Ryzen AI Halo vs Mac Studio: which is faster for LLMs?
The Mac Studio, on token generation. The Ryzen AI Max+ 395 (Strix Halo, sold in the Ryzen AI Halo box) uses LPDDR5X-8000 on a 256-bit bus, which is 256 GB/s. Apple lists 460 GB/s or 614 GB/s for the M5 Max Mac Studio and 1.2 TB/s for the M5 Ultra. Strix Halo wins on memory per dollar: 128GB, of which AMD says up to 96GB can be set as graphics memory. Both hold models that no single consumer GPU holds.
See Also
- Which Mac mini to Buy for Local LLMs: every config with its bandwidth and price
- Which Strix Halo Mini-PC to Buy: 256 GB/s boxes compared
- Best Local LLM for 64GB RAM: the models that fit this tier
- —n-cpu-moe Explained: llama.cpp MoE Offload Flags: run an MoE model across GPU and system RAM
- GGUF Quant Names Explained: read the real file size before you divide
- llama.cpp Flags Explained: the flags that change memory use and speed
- The Soldered Memory Trap: why you buy unified memory capacity once
Sources
- Apple Mac mini tech specs: M6 153 GB/s (16GB) and 170 GB/s (24GB, 32GB), M5 Pro 307 GB/s; read 2026-09-24
- Apple Mac Studio tech specs: M5 Max 460 GB/s and 614 GB/s, M5 Ultra 1.2 TB/s; read 2026-09-24
- NVIDIA DGX Spark: 128GB LPDDR5x, 273 GB/s; read 2026-09-24
- NVIDIA RTX Blackwell GPU architecture whitepaper: RTX 5090 28 Gbps GDDR7, 1,792 GB/s; read 2026-09-24
- NVIDIA Ampere GA102 whitepaper: RTX 3090 19.5 Gbps GDDR6X, 936 GB/s; read 2026-09-24
- NVIDIA RTX 5060 family page (128-bit GDDR7) and TechPowerUp RTX 5060 Ti PCIe x8 scaling review (PCIe 5.0 x8, 448 GB/s); read 2026-09-24
- AMD Ryzen AI Max+ 395: LPDDR5X-8000, 256-bit, 256 GB/s; AMD Variable Graphics Memory FAQ: up to 96GB as graphics memory; read 2026-09-24
- Threadripper memory channels: Phoronix TRX50 quad-channel review and AMD’s Threadripper PRO 9000 WX announcement (8 channels, DDR5-6400); read 2026-09-24
- PCIe rates: 16 GT/s (4.0) and 32 GT/s (5.0) per lane, 128b/130b encoding, per PCI-SIG; bandwidth values computed
- llama.cpp discussion #4167: Apple Silicon llama-bench table, Llama 2 7B Q4_0 3.56 GiB; community-reported; read 2026-09-24
- llama.cpp issue #4716: 7950X DDR5 vs 5900X DDR4 at 3,600 MT/s, Mixtral Q8_0; community-reported; read 2026-09-24
- Hugging Face API (
?blobs=true): unsloth/Qwen3.6-27B-GGUF Q4_K_M 16.82GB, unsloth/Qwen3.6-35B-A3B-GGUF UD-Q4_K_M 22.13GB, openai/gpt-oss-120b 65.25GB; model cards for Qwen3.6-35B-A3B (35B total, 3B activated) and gpt-oss-120b (117B, 5.1B active); Mistral: Mixtral of experts (12.9B per token); read 2026-09-24
Need OpenClaw fixed live?
Remote rescue sessions for gateway, auth, tunnel, VPS, and model access problems.
See Rescue Session