← All guides

Best Local LLM for Mac Studio (2026): gpt-oss 120B at 96GB+

The Mac Studio changed generation on 25 August 2026, and the answer to this question now depends on which of six memory tiers you buy. The M5 Max starts at $2,499 with 36GB and goes to 128GB. The M5 Ultra starts at $5,499 with 96GB and goes to 256GB now and 512GB in late October. From 96GB up, the model to run is gpt-oss 120B. Below that, it is a 27B-class machine. Nothing ships until 22 September, so every M5 speed figure here is derived from measured M3 Ultra and M4 Max results, and is labelled as such.

🛒 THE OUTGOING 128GB STUDIO, WHILE STOCK LASTS Mac Studio M4 Max 128GB · 546 GB/s Apple stopped selling it on 25 August 2026. It keeps 89% of the new M5 Max's bandwidth and runs gpt-oss 120B today, not on 22 September. Check current price on Amazon →

Bottom Line (September 2026)

  • From 96GB up, run gpt-oss 120B. 117B total parameters, 5.1B active, native MXFP4, about 65GB of weights. Owners of 96GB M3 Ultras report about 23 tok/s in everyday runtimes.
  • Below 96GB, it is a 27B machine. Qwen 3.6 27B at Q4 (~16GB) or Qwen 3.8 27B at Q4_K_M (~18GB). gpt-oss 120B does not fit in 64GB.
  • The lineup changed on 25 August 2026. M5 Max from $2,499 (36GB to 128GB, up to 614 GB/s). M5 Ultra from $5,499 (96GB to 512GB, 1.2 TB/s). M4 Max and M3 Ultra are discontinued.
  • The $4,000 step from 96GB to 256GB buys capacity, not speed. Same 1.2 TB/s. It only pays if you need 235B-class models.
  • Apple’s 4x claim is prefill, not generation. Token speed follows bandwidth: expect about 1.5x over M3 Ultra and about 1.1x over M4 Max.
  • Best local-LLM dollar right now: a discontinued M4 Max 128GB if a retailer still has one. It runs the same model at 89% of the new speed, today.

Ready to buy? See the tested hardware list with current prices.

Prices are Apple US prices, as of September 2026. Memory upgrade steps are not quoted here except the one Apple published; check the configurator for the rest.

The Lineup Changed on 25 August 2026

Most pages answering “best local LLM for Mac Studio” describe the M4 Max or the M3 Ultra. Apple sells neither. Here is what it sells.

ModelMemory optionsBandwidthStarting price
Mac Studio M5 Max (18-core CPU, 32-core GPU)36GB → 48GB → 64GB → 128GB460 GB/s$2,499 (36GB / 512GB)
Mac Studio M5 Max (18-core CPU, 40-core GPU)same614 GB/scheck configurator
Mac Studio M5 Ultra (30-core CPU, 64-core GPU)96GB → 256GB → 512GB (late October)1.2 TB/s$5,499 (96GB / 1TB)
Mac Studio M4 Max (discontinued)up to 128GB546 GB/sretail clearance, check listings
Mac Studio M3 Ultra (discontinued)96GB at the end819 GB/sretail clearance, check listings

Apple announced the refresh on 25 August 2026. Machines arrive on 22 September. The 256GB M5 Ultra adds $4,000 to the 96GB price, so it starts at $9,499. The 512GB tier ships in late October and has no published price yet.

One detail decides the M5 Max choice. The 32-core GPU version runs memory at 460 GB/s. The 40-core GPU version runs it at 614 GB/s. Token generation is bandwidth-bound, so the second one is a third faster at the same model. If you buy an M5 Max for local AI, buy the 40-core GPU.

The Number Nobody States

macOS reserves roughly 25% of unified memory for the system and hands the GPU about 75%. That single fact places each Mac Studio in a model class.

Mac Studio memoryRoughly usable for modelsWhat fits
36GB~27GB27B at Q4 to Q6
48GB~36GB27B at Q8, or 35B with long context
64GB~48GB35B comfortably; dense 70B at Q4 (~40GB) with almost no context
96GB~72GBgpt-oss 120B (~65GB) with usable context
128GB~96GBgpt-oss 120B at full 128K context, plus a second model
256GB~192GB235B-class MoE at Q4; DeepSeek V4 Flash (~80GB Q4) with room to spare

You can raise the allocation with iogpu.wired_limit_mb. It recovers a few gigabytes. It does not move you up a tier.

Picks by Memory Tier

36GB to 64GB M5 Max — the 27B tier

ModelQuantSizeWhy
Qwen 3.6 27BQ4_K_M~16GBBest all-round pick. Fits every tier with a real context window.
Qwen 3.8 27BQ4_K_M~18GBNewer, Apache 2.0, includes a 931MB vision encoder.
Llama 3.3 70BQ4_K_M~40GB64GB only. Fits on paper, leaves ~8GB for context.

At 64GB the dense 70B is the trap. It fits, and then every token it writes is a 40GB read, which at 614 GB/s caps below 15 tok/s before context costs anything. A 27B at Q8 is faster and often no worse at agent work.

96GB and 128GB — gpt-oss 120B

Our pick: gpt-oss 120B. It is the only frontier-adjacent open model that fits 96GB with usable context, and it is the most reliable tool-caller in the open-weight set, which matters more than benchmark score when it runs OpenClaw unattended. On 96GB M3 Ultra machines, owners report about 23 tok/s with a 2.3s time to first token in everyday runtimes, and tuned MLX setups reach far higher.

At 128GB you gain two things: the full 128K context window without a memory squeeze, and headroom to keep a 9B model resident for routing and embeddings. On a 128GB M4 Max at 546 GB/s, the same model measured 14–20 tok/s at Q6. Scale that by 614 ÷ 546 and the M5 Max 128GB should land at roughly 16–22 tok/s. That is a derived figure and it is marked as one.

If you own an M5 Max 128GB and want a coding model instead, Laguna S 2.1 (118B total, 8B active, 1M context) runs in the same footprint. See the Laguna S 2.1 setup guide.

256GB M5 Ultra — the 235B tier

Qwen3-VL 235B at Q4_K_M is the daily driver on a 256GB Mac Studio. On the M3 Ultra it ran at about 30 tok/s, it sees images, and it leaves room for a second model. GLM-4.7 (358B) at Q3 is the intelligence ceiling at about 15 tok/s on M3 Ultra, which is too slow to chat with and right for a batch agent. Multiply both by about 1.5 for the M5 Ultra once real numbers exist. The 256GB Mac Studio model guide has the full table.

What the $4,000 Buys, and What It Does Not

The 96GB and 256GB M5 Ultra share the same 1.2 TB/s bus. So gpt-oss 120B runs at the same speed on both. The extra $4,000 buys the ability to load a model that does not fit in 96GB. If you are not going to run a 235B-class model, that money buys nothing you can measure.

The step that does change speed is Max to Ultra. The M5 Ultra moves data at about 2x the rate of the 40-core M5 Max (1.2 TB/s against 614 GB/s). A 128GB M5 Max and a 96GB M5 Ultra both run gpt-oss 120B; the Ultra runs it roughly twice as fast and costs about $3,000 more before the Max’s memory upgrade is counted.

How to Predict Your Own Token Speed

Rough ceiling in tokens/sec = memory bandwidth ÷ bytes read per token. For a dense model, that is the whole file. For a mixture-of-experts model like gpt-oss 120B, it is the active experts plus attention, which is why a 65GB model generates 23 tok/s and not 12. Expect 60% to 70% of the ceiling in practice.

Apple’s launch claim of “up to 4x faster LLM prompt processing” on M5 Ultra against M3 Ultra is real and it is about prefill, which is compute-bound and benefits from the new GPU Neural Accelerators. Generation is bandwidth-bound, and bandwidth rose 50% (819 GB/s to 1.2 TB/s) on the Ultra and 12% (546 to 614 GB/s) on the Max. Time to first token on a long prompt will improve a lot. Reading speed will improve by about half.

The Honest Recommendation

If you can find a 128GB M4 Max Mac Studio on clearance, buy it. Apple stopped selling it on 25 August 2026. It keeps 546 GB/s, which is 89% of the new M5 Max, and it runs gpt-oss 120B today instead of after 22 September. Retailers still list it — check the current Mac Studio M4 Max 128GB price and check the memory size on the listing, because the 36GB and 64GB units share the same name.

If you are buying new and want gpt-oss 120B, the cheapest new path is the M5 Max with the 40-core GPU and 128GB. Apple has not published that configuration’s price on its newsroom, so check the configurator. Do not buy the 36GB base for local AI; it is a 27B machine at a 120B price bracket.

If you want speed and can spend $5,499, the M5 Ultra 96GB is the pick. Same model as the 128GB Max, about twice the token rate. Skip the 256GB step unless you have a named 235B-class model you intend to run.

If you need 512GB, wait. Apple says late October, with no price. Our wait-for-M5-Ultra guide covers the timing.

See Also

Need OpenClaw fixed live?

Remote rescue sessions for gateway, auth, tunnel, VPS, and model access problems.

See Rescue Session

Read next

Which Mac Studio Should You Buy for Local LLMs? (2026)
The exact Mac Studio configuration to order for local LLMs, as of September 2026. Eight M5 Max and M5 Ultra configs from $2,499 to $10,799, priced per GB of unified memory, with one pick per budget. The 256GB tier now has a concrete reason: MiMo-V2.6-Flash needs about 165GB and does not fit a 128GB Mac.
Best Local LLM for Mac Studio M5 Ultra (2026): 256GB
Apple announced the M5 Ultra Mac Studio on August 25, 2026 with 1.2TB/s bandwidth and a 256GB option — the first new 256GB Mac since Apple pulled the tier in May. What fits, projected tok/s, the $4,000 memory tax, and why you should wait for real benchmarks.
Is 96GB of VRAM Enough for Local AI in 2026?
96GB is the first tier where a dense 70B runs at its full 128K window: 42.5GB of Q4 weights plus exactly 40 GiB of FP16 KV cache is 82.5GB, and it fits. What 96GB unlocks, what it still cannot hold, and what the one card that has it costs in 2026.
Strix Halo vs Mac Studio M4 Max 128GB for Local LLMs
Compare AMD Strix Halo (Ryzen AI Max+ 395) and Mac Studio M4 Max 128GB for local LLMs: 256 vs 546 GB/s bandwidth, decode speed, ROCm vs MLX, 2026 prices.