← All guides

Best Local LLM by RAM (July 2026): 8GB to 128GB Picks

Your RAM decides which local LLM you can run. This hub maps every common tier from 8GB to 128GB to the model that fits in July 2026. Quick answer: run Qwen 3.6 27B at 24-32GB, Laguna XS 2.1 for agentic coding at 48-64GB, and gpt-oss 120B at 64-128GB. Below each tier links to a full guide with quant sizes and OpenClaw tool-call notes.

Need help picking the right model for your hardware?

See our AI training options. We'll match your RAM to the right model and quant in 30 minutes.

Pick Your RAM Tier (July 2026)

Your RAMBest PickBest For OpenClawDetailed Guide
8 GBQwen 3.5 4B (Q5_K_M)Not recommended — use cloud8GB guide →
16 GBQwen 3.5 9B (Q5_K_M)gpt-oss 20B (Q4)16GB guide →
24 GBQwen 3.6 27B (Q4_K_M) ← NEWgpt-oss 20B (Q5)24GB guide →
32 GBQwen 3.6 27B (Q6_K)Qwen 3.6 27B / gpt-oss 20B (Q8)32GB guide →
48 GBLaguna XS 2.1 (Q8) ← coding / Qwen 3.6 35B-A3B (Q5)Qwen 3.6 27B (Q8)48GB guide →
64 GBgpt-oss 120B (Q4_K_M) / Laguna XS 2.1 for codinggpt-oss 120B / Mistral Small 4 (119B-A6B)64GB guide →
96 GBQwen 3.5 122B-A10B (Q4_K_M)gpt-oss 120B (Q5)96GB guide →
128 GBgpt-oss 120B (Q6_K)gpt-oss 120B (Q8)128GB guide →

If you came from a Reddit-style model search, use the short versions first: Best local LLM Reddit picks for OpenClaw, Reddit’s favorite local LLM for OpenClaw, Best OpenClaw model Reddit users recommend, 32GB Reddit shortlist, 64GB Reddit shortlist, or 128GB Reddit shortlist. Those pages answer the “what does Reddit recommend?” phrasing directly, then route back here for the full RAM table.

What Changed by July 2026

The local LLM landscape moved fast between February and July 2026:

  • Laguna XS 2.1 (Poolside) — 33B/3B MoE with a 256K context. It fits in about 36GB at Q8 and is now the best agentic-coding pick for the 48-64GB tiers. It is in the Ollama library.
  • DeepSeek V4 Flash — Runs on Ollama cloud only (use the :cloud command). Local V4 support is still experimental forks, so treat it as cloud-first at every consumer RAM tier.
  • Qwen 3.6 27B (April 22) — Dense 27B that outperforms the 397B Qwen 3.5 MoE on agentic coding (77.2 vs 76.x on SWE-Bench Verified). Still the default for 24-32GB tiers.
  • GLM-5.1 (April 7) — 744B MoE from Z.ai. Cloud-only. (Earlier guides citing “GLM-5.1 32B” were referring to the older GLM-4 line, not 5.1.)
  • Mistral Small 4 (March 16) — 119B-A6B MoE that fits at Q4 in about 60GB. Replaces Mistral Large 123B.
  • Qwen 3.5 small series (March 2) — 0.8B / 2B / 4B / 9B variants. The 9B is the new 16GB tier pick.
  • Qwen 3.5 medium (February 24) — 27B dense, 35B-A3B MoE, 122B-A10B MoE. The 35B-A3B MoE is excellent at 48GB.
  • Llama 3.3 70B — Still works, no longer the default. The Qwen and gpt-oss families have caught up at smaller sizes.

How to Use This Guide

Step 1: Find your usable RAM, not your installed RAM. On Mac, the OS reserves 4-6GB. On Windows or Linux with an NVIDIA GPU, the relevant number is VRAM (the GPU’s onboard memory), not system RAM.

Step 2: Subtract context overhead. A 32K context window costs roughly 4-6GB. A 128K window costs 16-24GB. Model weights are not the only thing that has to fit.

Step 3: Pick the highest-quality quant that leaves headroom. Q5_K_M is the sweet spot. Q4_K_M is the standard. Below Q3 starts to hurt tool calling, which kills agent runs.

OpenClaw Tool-Calling Reality Check (July 2026)

Most local LLM guides talk about benchmark scores. For OpenClaw, only one metric matters: does the model emit valid JSON when asked to call a tool, hundreds of times in a row, without drift?

Models that pass this filter today:

  • gpt-oss 20B — cleanest tool-call JSON in production, this is the safe default
  • gpt-oss 120B — same family, scaled up
  • Qwen 3.6 27B — fixed the tool-calling regressions from 3.5
  • Qwen 3.6 35B-A3B (MoE) — fast inference with reliable tools
  • Llama 3.3 70B — still fine for tool calls
  • Mistral Small 4 (119B-A6B) — works, but heavier than gpt-oss

Models to avoid for OpenClaw right now:

  • Qwen 3.5 27B — known broken tool-calling in Ollama (GitHub issue #14493)
  • Anything under 7B — too unreliable for autonomous loops
  • Most fine-tunes of base models

Quantization Cheat Sheet

QuantBits/weightQualityWhen to use
Q8_08Near-FP16When you have 2x the model size in RAM
Q5_K_M~5.5Indistinguishable from Q8Best quality-to-size ratio
Q4_K_M~4.5Loses 1-3% on benchmarksStandard pick when RAM is tight
IQ3_XS~3.3Noticeable degradation, MoE-friendlySqueeze a bigger model into too-little RAM
Q2_K~2.6Significantly degradedLast resort, breaks tool calling

Can Your RAM Run It? (exact-answer guides)

See Also

Apple Silicon chip guides:

Need OpenClaw fixed live?

Remote rescue sessions for gateway, auth, tunnel, VPS, and model access problems.

See Rescue Session

Read next

Best Local LLMs for 8GB RAM (April 2026): Qwen 3.5 Small Series
The best local LLMs that fit in 8GB RAM or 8GB VRAM. April 2026 picks: Qwen 3.5 4B, Qwen 3.5 9B (squeeze), gpt-oss 20B at IQ2, with quants and OpenClaw notes.
Best Local LLM for Mac Studio M2 Ultra (2026): 64/128/192 GB Unified
Best local LLM for the Mac Studio M2 Ultra. April 2026 picks for 64GB, 128GB, 192GB variants. gpt-oss 120B, Mistral Small 4 (119B-A6B), Llama 3.3 70B Q8, and quad-model OpenClaw setups.
Best Local LLMs for 128GB RAM (July 2026): Llama 4 Maverick, gpt-oss 120B & Laguna XS 2.1
Best local LLMs for 128GB RAM in July 2026. Llama 4 Maverick (400B MoE, ~95GB Q4), gpt-oss 120B at Q6, Laguna XS 2.1 (agentic coding, Q8 + huge context), Llama 4 Scout (10M context), DeepSeek V4 Flash via Ollama cloud. Mac Studio M4 Max territory.
Best Local LLM for MacBook Pro M4 Max (July 2026): 36 to 128GB Picks
Best local LLM for the MacBook Pro M4 Max, updated July 2026. Tier picks: 36GB Qwen 3.6 27B Q6, 64GB Llama 3.3 70B Q5, 128GB Mistral Small 4. Coding pick: Laguna XS 2.1.