Best Local LLM for Mac Studio (2026): gpt-oss 120B at 96GB+
The Mac Studio changed generation on 25 August 2026, and the answer to this question now depends on which of six memory tiers you buy. The M5 Max starts at $2,499 with 36GB and goes to 128GB. The M5 Ultra starts at $5,499 with 96GB and goes to 256GB now and 512GB in late October. From 96GB up, the model to run is gpt-oss 120B. Below that, it is a 27B-class machine. Nothing ships until 22 September, so every M5 speed figure here is derived from measured M3 Ultra and M4 Max results, and is labelled as such.
Bottom Line (September 2026)
- From 96GB up, run gpt-oss 120B. 117B total parameters, 5.1B active, native MXFP4, about 65GB of weights. Owners of 96GB M3 Ultras report about 23 tok/s in everyday runtimes.
- Below 96GB, it is a 27B machine. Qwen 3.6 27B at Q4 (~16GB) or Qwen 3.8 27B at Q4_K_M (~18GB). gpt-oss 120B does not fit in 64GB.
- The lineup changed on 25 August 2026. M5 Max from $2,499 (36GB to 128GB, up to 614 GB/s). M5 Ultra from $5,499 (96GB to 512GB, 1.2 TB/s). M4 Max and M3 Ultra are discontinued.
- The $4,000 step from 96GB to 256GB buys capacity, not speed. Same 1.2 TB/s. It only pays if you need 235B-class models.
- Apple’s 4x claim is prefill, not generation. Token speed follows bandwidth: expect about 1.5x over M3 Ultra and about 1.1x over M4 Max.
- Best local-LLM dollar right now: a discontinued M4 Max 128GB if a retailer still has one. It runs the same model at 89% of the new speed, today.
Ready to buy? See the tested hardware list with current prices.
Prices are Apple US prices, as of September 2026. Memory upgrade steps are not quoted here except the one Apple published; check the configurator for the rest.
The Lineup Changed on 25 August 2026
Most pages answering “best local LLM for Mac Studio” describe the M4 Max or the M3 Ultra. Apple sells neither. Here is what it sells.
| Model | Memory options | Bandwidth | Starting price |
|---|---|---|---|
| Mac Studio M5 Max (18-core CPU, 32-core GPU) | 36GB → 48GB → 64GB → 128GB | 460 GB/s | $2,499 (36GB / 512GB) |
| Mac Studio M5 Max (18-core CPU, 40-core GPU) | same | 614 GB/s | check configurator |
| Mac Studio M5 Ultra (30-core CPU, 64-core GPU) | 96GB → 256GB → 512GB (late October) | 1.2 TB/s | $5,499 (96GB / 1TB) |
| Mac Studio M4 Max (discontinued) | up to 128GB | 546 GB/s | retail clearance, check listings |
| Mac Studio M3 Ultra (discontinued) | 96GB at the end | 819 GB/s | retail clearance, check listings |
Apple announced the refresh on 25 August 2026. Machines arrive on 22 September. The 256GB M5 Ultra adds $4,000 to the 96GB price, so it starts at $9,499. The 512GB tier ships in late October and has no published price yet.
One detail decides the M5 Max choice. The 32-core GPU version runs memory at 460 GB/s. The 40-core GPU version runs it at 614 GB/s. Token generation is bandwidth-bound, so the second one is a third faster at the same model. If you buy an M5 Max for local AI, buy the 40-core GPU.
The Number Nobody States
macOS reserves roughly 25% of unified memory for the system and hands the GPU about 75%. That single fact places each Mac Studio in a model class.
| Mac Studio memory | Roughly usable for models | What fits |
|---|---|---|
| 36GB | ~27GB | 27B at Q4 to Q6 |
| 48GB | ~36GB | 27B at Q8, or 35B with long context |
| 64GB | ~48GB | 35B comfortably; dense 70B at Q4 (~40GB) with almost no context |
| 96GB | ~72GB | gpt-oss 120B (~65GB) with usable context |
| 128GB | ~96GB | gpt-oss 120B at full 128K context, plus a second model |
| 256GB | ~192GB | 235B-class MoE at Q4; DeepSeek V4 Flash (~80GB Q4) with room to spare |
You can raise the allocation with iogpu.wired_limit_mb. It recovers a few gigabytes. It does not move you up a tier.
Picks by Memory Tier
36GB to 64GB M5 Max — the 27B tier
| Model | Quant | Size | Why |
|---|---|---|---|
| Qwen 3.6 27B | Q4_K_M | ~16GB | Best all-round pick. Fits every tier with a real context window. |
| Qwen 3.8 27B | Q4_K_M | ~18GB | Newer, Apache 2.0, includes a 931MB vision encoder. |
| Llama 3.3 70B | Q4_K_M | ~40GB | 64GB only. Fits on paper, leaves ~8GB for context. |
At 64GB the dense 70B is the trap. It fits, and then every token it writes is a 40GB read, which at 614 GB/s caps below 15 tok/s before context costs anything. A 27B at Q8 is faster and often no worse at agent work.
96GB and 128GB — gpt-oss 120B
Our pick: gpt-oss 120B. It is the only frontier-adjacent open model that fits 96GB with usable context, and it is the most reliable tool-caller in the open-weight set, which matters more than benchmark score when it runs OpenClaw unattended. On 96GB M3 Ultra machines, owners report about 23 tok/s with a 2.3s time to first token in everyday runtimes, and tuned MLX setups reach far higher.
At 128GB you gain two things: the full 128K context window without a memory squeeze, and headroom to keep a 9B model resident for routing and embeddings. On a 128GB M4 Max at 546 GB/s, the same model measured 14–20 tok/s at Q6. Scale that by 614 ÷ 546 and the M5 Max 128GB should land at roughly 16–22 tok/s. That is a derived figure and it is marked as one.
If you own an M5 Max 128GB and want a coding model instead, Laguna S 2.1 (118B total, 8B active, 1M context) runs in the same footprint. See the Laguna S 2.1 setup guide.
256GB M5 Ultra — the 235B tier
Qwen3-VL 235B at Q4_K_M is the daily driver on a 256GB Mac Studio. On the M3 Ultra it ran at about 30 tok/s, it sees images, and it leaves room for a second model. GLM-4.7 (358B) at Q3 is the intelligence ceiling at about 15 tok/s on M3 Ultra, which is too slow to chat with and right for a batch agent. Multiply both by about 1.5 for the M5 Ultra once real numbers exist. The 256GB Mac Studio model guide has the full table.
What the $4,000 Buys, and What It Does Not
The 96GB and 256GB M5 Ultra share the same 1.2 TB/s bus. So gpt-oss 120B runs at the same speed on both. The extra $4,000 buys the ability to load a model that does not fit in 96GB. If you are not going to run a 235B-class model, that money buys nothing you can measure.
The step that does change speed is Max to Ultra. The M5 Ultra moves data at about 2x the rate of the 40-core M5 Max (1.2 TB/s against 614 GB/s). A 128GB M5 Max and a 96GB M5 Ultra both run gpt-oss 120B; the Ultra runs it roughly twice as fast and costs about $3,000 more before the Max’s memory upgrade is counted.
How to Predict Your Own Token Speed
Rough ceiling in tokens/sec = memory bandwidth ÷ bytes read per token. For a dense model, that is the whole file. For a mixture-of-experts model like gpt-oss 120B, it is the active experts plus attention, which is why a 65GB model generates 23 tok/s and not 12. Expect 60% to 70% of the ceiling in practice.
Apple’s launch claim of “up to 4x faster LLM prompt processing” on M5 Ultra against M3 Ultra is real and it is about prefill, which is compute-bound and benefits from the new GPU Neural Accelerators. Generation is bandwidth-bound, and bandwidth rose 50% (819 GB/s to 1.2 TB/s) on the Ultra and 12% (546 to 614 GB/s) on the Max. Time to first token on a long prompt will improve a lot. Reading speed will improve by about half.
The Honest Recommendation
If you can find a 128GB M4 Max Mac Studio on clearance, buy it. Apple stopped selling it on 25 August 2026. It keeps 546 GB/s, which is 89% of the new M5 Max, and it runs gpt-oss 120B today instead of after 22 September. Retailers still list it — check the current Mac Studio M4 Max 128GB price and check the memory size on the listing, because the 36GB and 64GB units share the same name.
If you are buying new and want gpt-oss 120B, the cheapest new path is the M5 Max with the 40-core GPU and 128GB. Apple has not published that configuration’s price on its newsroom, so check the configurator. Do not buy the 36GB base for local AI; it is a 27B machine at a 120B price bracket.
If you want speed and can spend $5,499, the M5 Ultra 96GB is the pick. Same model as the 128GB Max, about twice the token rate. Skip the 256GB step unless you have a named 235B-class model you intend to run.
If you need 512GB, wait. Apple says late October, with no price. Our wait-for-M5-Ultra guide covers the timing.
See Also
- Which Mac Studio Should You Buy for Local LLMs? — every M5 configuration priced from Apple’s configurator, and the one to order
- Mac mini vs Mac Studio for Local LLMs — whether you need the Studio at all
- Best Local LLM for Mac Studio M5 Ultra 256GB — the top tier in depth
- Best Local LLM for Mac Studio M4 — the outgoing M4 Max Studio
- Best Models for the 256GB Mac Studio — the 235B-class table
- Mac Studio vs RTX Workstation for Local LLMs — Apple bandwidth vs CUDA
- Best Local LLM for Mac mini — one step down, same 75% rule
- Strix Halo vs Mac Studio M4 Max at 128GB — the $2,000 alternative at the same memory
Need OpenClaw fixed live?
Remote rescue sessions for gateway, auth, tunnel, VPS, and model access problems.
See Rescue Session