← All guides

Best Local LLMs for 48GB RAM (July 2026): Qwen 3.6 27B Q8 + Laguna XS 2.1 for Coding

48GB is a solid tier in July 2026. Qwen 3.6 27B at Q8 remains the best overall pick. New this month: Poolside's Laguna XS 2.1 (33B/3B MoE, 256K context) fits at about 36GB at Q8 and is the best agentic coding model you can run at this tier — cap the context to keep it inside 48GB. Note: Llama 4 Scout needs ~58-60GB and does NOT fit 48GB — you need 64GB for Scout.

Running 8-hour OpenClaw agents on M3 Max?

See our AI training options. We'll dial in dual-model routing + context strategy + launchd for unattended overnight runs.

Apple Mac for 48GB RAM local AI on Amazon
🛒 BEST MAC FOR 48GB RAM Apple Mac Studio · 48GB+ unified memory 48GB unified memory runs 70B-class models and multi-model setups — the recommended Mac for serious local AI. Check current price on Amazon →
Updated July 22, 2026 — Laguna XS 2.1 reaches this tier
  • Laguna XS 2.1 (Poolside, 33B/3B MoE) — ~36GB at Q8, 256K context, best agentic coding at 48GB; cap the context or use INT4 weights
  • Gemma 4 26B-A4B (Google, June 3) — 26B MoE, ~15GB at Q4, 45-65 tok/sec, best fast secondary model
  • Devstral Small 24B (Mistral) — dedicated coding model, ~14.5GB at Q4, strong HumanEval
  • Llama 4 Scout (10M context) needs ~58GB — does not fit 48GB. Need 64GB for Scout.

Bottom Line (July 2026)

  • Best overall pick: Qwen 3.6 27B at Q8_0 (near-FP16 quality, 30GB footprint)
  • Best agentic coding: Laguna XS 2.1 (33B/3B MoE) at Q8 — ~36GB, 256K context, in Ollama
  • Best for fast inference: Qwen 3.6 35B-A3B MoE at Q6_K (40-60 tok/sec)
  • Best for OpenClaw production: Dual — gpt-oss 20B Q8 + Qwen 3.6 27B Q5
  • Best new lightweight model: Gemma 4 26B-A4B — ~15GB, 45-65 tok/sec, great second model
  • Best small coding model: Devstral Small 24B — fits at 14.5GB, leaves room for a second model

Top Picks for 48GB RAM

1. Qwen 3.6 27B (Q8_0) — best general-purpose at premium quality

Q8_0 of the April 22 release uses about 30GB and gives near-FP16 quality. The “ship it forever” pick at this tier. Speed: 25-40 tok/sec on M3 Max.

ollama pull qwen3.6:27b-q8_0
openclaw config set agents.defaults.models.chat ollama/qwen3.6:27b-q8_0

2. Laguna XS 2.1 (33B/3B MoE) at Q8 — best agentic coding [New July 2026]

Poolside’s July 2026 agentic coding model. 33B total parameters with 3B active per token (MoE), a 256K context window, and up to 32K output tokens. At Q8/FP8 the weights use about 33-36GB. It gains +5.4% on SWE-bench Multilingual over Laguna XS.2 and is built for tool calling and long terminal runs.

48GB is the tight end for this model. After macOS overhead you have about 38-40GB, so the Q8 weights leave only a few GB for the KV cache. Cap the context at 32-64K, or pull the INT4/NVFP4 weights if you need the full 256K window.

ollama pull laguna-xs-2.1
openclaw config set agents.defaults.models.agent ollama/laguna-xs-2.1
openclaw config set agents.defaults.context 65536
openclaw run --agent "Fix the failing tests and open a PR"

Weights ship in BF16, FP8, NVFP4, and INT4 on Hugging Face. The FP8 KV cache keeps memory flat during long agent loops, which matters more at 48GB than at 64GB.

3. Qwen 3.6 35B-A3B (Q6_K) — fastest at this tier

The Mixture-of-Experts variant of Qwen 3.6 at Q6_K uses about 30GB. 35B total parameters with 3B active per token = 8B-class inference speed with 35B-class knowledge. The right pick if you do many short interactions.

ollama pull qwen3.6:35b-q6_K
openclaw config set agents.defaults.models.chat ollama/qwen3.6:35b-q6_K

4. Dual-Model OpenClaw Setup (the 48GB advantage)

Keep two specialized models loaded for instant routing:

# gpt-oss 20B Q8 for autonomous agent runs (cleanest tool calls) — 22GB
# Qwen 3.6 27B Q5 for general chat (premium reasoning) — 20GB

openclaw config set agents.defaults.models.chat ollama/qwen3.6:27b-q5_K_M
openclaw config set agents.defaults.models.agent ollama/gpt-oss:20b-q8_0
openclaw config set agents.defaults.keep_alive 30m

# Verify
openclaw models status

This routing pattern is unique to 48GB+ tiers. Below this, model swap latency hurts.

5. Nemotron Cascade 2 30B (Q8_0) — premium structured output

NVIDIA’s late-March 2026 release at Q8 uses about 32GB. Strongest open model for JSON output and structured generation at this RAM tier.

ollama pull nemotron-cascade-2:30b-q8_0

6. Mistral Small 4 (119B-A6B MoE, IQ3_XS) — squeeze for the new Mistral

Mistral’s March 16, 2026 release replaces Mistral Large 123B. The 119B-A6B MoE at IQ3_XS uses about 38GB. 6B active params per token = fast inference. Quality is degraded at IQ3 but still useful.

ollama pull mistral-small-4:iq3_xs

What Fits in 48GB

ModelQuantRAM UsedTool Calling
Qwen 3.6 27BQ8_0~33 GBExcellent
Qwen 3.6 35B-A3BQ6_K~33 GBExcellent
Laguna XS 2.1 33B-A3BQ8/FP8~36 GB (cap context)Excellent
Nemotron Cascade 2 30BQ8_0~34 GBGood
Mistral Small 4 119B-A6BIQ3_XS~40 GBGood
Qwen 3.5 122B-A10BIQ3_XS~42 GBFair (Ollama bug)
gpt-oss 20B + Qwen 3.6 27B Q5 (dual)Q8 + Q5~42 GBExcellent

Common Mistakes at 48GB

  1. Defaulting to Llama 3.3 70B at Q3 because “bigger is better”. Qwen 3.6 27B at Q8 now outperforms Llama 3.3 70B Q4 on most agentic tasks.
  2. Running Q8 of a 27B with 256K context. KV cache eats 30GB+ on top of the model. Cap at 64K for Q8.
  3. Forgetting the OS uses RAM too. macOS Sonoma/Sequoia uses 6-10GB during normal use. Treat 48GB as 38-40GB available.
  4. Running Laguna XS 2.1 at Q8 with the full 256K context. The weights already take ~36GB. The KV cache then pushes you into swap. Cap the context at 32-64K, or use the INT4 weights.
  5. Picking Qwen 3.5 122B-A10B for OpenClaw. Tool calling bug affects this MoE too. Use Qwen 3.6 27B/35B-A3B instead.

Hardware That Actually Hits 48GB

  • M3 Max MacBook Pro (48GB) — best laptop pick
  • M4 Max MacBook Pro (48GB)
  • Mac Studio M2 Max (64GB) — close enough, gives headroom
  • NVIDIA RTX A6000 48GB — workstation, single card
  • 2x RTX 3090 24GB — 48GB total VRAM (Linux setup, complex)

See Also

Need OpenClaw fixed live?

Remote rescue sessions for gateway, auth, tunnel, VPS, and model access problems.

See Rescue Session

Read next

Best Local LLMs for 64GB RAM (July 2026): Llama 4 Scout, gpt-oss 120B, DeepSeek V4 Flash & Laguna XS 2.1
Best local LLMs for 64GB RAM in July 2026. Llama 4 Scout (10M context, ~58GB Q4), gpt-oss 120B at Q4, DeepSeek V4 Flash (284B MoE, Ollama cloud), Laguna XS 2.1 (agentic coding, 33B-A3B, ~36GB Q8). Also: Mistral Small 4, Qwen 3.6 35B Q8.
Best Local LLM for 32GB RAM (July 2026): Qwen 3.6 27B + What to Avoid
Best local LLM for 32GB RAM right now: Qwen 3.6 27B Q6_K (~22GB, fits with headroom). Full tested list — what fits, what barely fits, what to avoid — plus exact Ollama commands and tok/sec numbers.
Best Local LLMs for 24GB RAM (April 2026): Qwen 3.6 27B Headlines
Best local LLMs for 24GB RAM in April 2026. Qwen 3.6 27B (released Apr 22) is the new headline pick — outperforms 397B MoE models on agentic coding. Plus gpt-oss 20B, Qwen 3.5 9B at Q8.
Best Local LLMs for 96GB RAM (June 2026): Llama 4 Scout, DeepSeek V4 Flash & gpt-oss 120B Q5
Best local LLMs for 96GB RAM in June 2026. Llama 4 Scout (10M context, ~58GB Q4), DeepSeek V4 Flash (~80GB Q4), gpt-oss 120B at Q5 (~80GB), Qwen 3.5 122B-A10B, Mistral Small 4 at Q5. Mac Studio M3 Ultra territory.