← All guides

DGX Spark vs Strix Halo for Local LLMs: $4,699 vs $2,600 128GB Boxes

Both boxes give you 128GB of unified memory for local LLMs. The DGX Spark buys you CUDA and clustering for $4,699. A Strix Halo mini PC starts around $2,600 as of August 2026. Single-user speed is close to a tie.

Before buying a 128GB box

Run your target model through the estimator first. 128GB of capacity does not fix 256–273 GB/s of bandwidth: dense 70B is slow on both boxes, and MoE models are the sweet spot.

Check the fit in the estimator
🎮 THE BOXES IN THIS COMPARISON

The DGX Spark and the ASUS Ascent GX10 share the same GB10 chip and 128 GB of memory. The AMD Ryzen AI Max+ 395 box is the cheaper 128 GB path. Prices move monthly in the shortage — check the live listing before deciding.

Short answer

Choose a Strix Halo box if your priority is:

  • lowest cost per GB of unified memory
  • Vulkan-based runtimes like llama.cpp and LM Studio
  • a general-purpose Windows or Linux PC that also runs LLMs
  • saving roughly $2,000 versus the Spark at the bottom of the range

Choose the DGX Spark if your priority is:

  • CUDA and the NVIDIA software stack out of the box
  • clustering two units over 200GbE for 200B–405B models
  • FP4 model support and NVIDIA’s tuned containers
  • matching the stack your cloud GPUs run

For single-user inference, speed does not decide this. Software and price do.

Specs that matter

BoxMemoryBandwidthComputePrice (Aug 2026)
NVIDIA DGX Spark (GB10)128GB LPDDR5x273 GB/s~1 PFLOP FP4$4,699
Strix Halo (Ryzen AI Max+ 395)128GB LPDDR5x256 GB/sRDNA 3.5 iGPUfrom ~$2,600

NVIDIA lists the DGX Spark with 128GB LPDDR5x at 273 GB/s, about 1 petaFLOP of FP4 compute, and built-in ConnectX-7 200GbE networking. The box is spec’d for 240W power delivery. AMD’s Ryzen AI Max+ 395 uses quad-channel LPDDR5x-8000 for 256 GB/s, and up to 96GB is assignable to the GPU.

The price story, as of August 2026

Both boxes got more expensive. NVIDIA raised the Spark from $3,999 to $4,699 in February 2026 and cited memory supply, per Tom’s Hardware and NVIDIA’s developer forum. Strix Halo boxes floored around $2,000 pre-shortage; a July 31 sweep of 21 Amazon 128GB listings found the cheapest near $2,600, with mid-range units like the GMKtec EVO-X3 at $3,500 and up. The Register lists the HP Z2 Mini G1a at $2,949 as a nameable mid-range option.

So the real gap is about $2,000 only if you buy the cheapest Strix Halo listing. Against a $3,500 mid-range box, the Spark’s premium shrinks to about $1,200.

Performance: near parity, with one big asterisk

The Register benchmarked both platforms. The text-level findings: single-batch inference is close to parity, AMD is sometimes ahead on Vulkan runtimes, and both boxes struggle with dense 70B models. Community results put dense 70B decode around 2.6 tok/s on either box. That is a demo, not a daily driver.

Both boxes shine on MoE models. A Qwen3-30B-A3B-class model activates ~3B parameters per token, so 256–273 GB/s of bandwidth is enough for comfortable speeds while the 128GB pool holds the full weights. If you buy either box, plan your model list around MoE. Our 128GB RAM model guide covers the current picks.

The Spark’s unique trick is clustering. NVIDIA states two Sparks link over the built-in ConnectX-7 200GbE and handle models up to 405B parameters. Strix Halo has no equivalent turnkey path.

Software: CUDA convenience vs ROCm progress

The Spark runs NVIDIA’s stack natively: CUDA, TensorRT-LLM, tuned NIM containers. Things mostly work on day one.

ROCm on Strix Halo is workable but rougher. In the Register’s testing, most PyTorch scripts ran unmodified, but vLLM and FlashAttention-2 needed manual compilation. If you live in llama.cpp and LM Studio on Vulkan, you will barely notice. If you want vLLM serving or exotic runtimes, budget tinkering time.

Decision table

Your situationBetter default
Budget-first, LM Studio / llama.cpp userStrix Halo
You want NVIDIA’s stack end to endDGX Spark
Planning to scale to 200B+ models laterDGX Spark (two clustered)
You also want a daily Windows PCStrix Halo
vLLM or PyTorch-heavy workflowsDGX Spark, or accept ROCm build work
Dense 70B is your main modelNeither — see a 24–32GB GPU plus quantization, or a Mac

Final recommendation

For most OpenClaw users who want 128GB on a budget: Strix Halo, at the bottom of the price range, running MoE models on Vulkan.

Pay for the DGX Spark when the CUDA stack or the two-box clustering path has concrete value for you. And re-check both prices the week you buy. As of August 2026 the memory shortage moves them monthly.

Next steps

Sources

Need OpenClaw fixed live?

Remote rescue sessions for gateway, auth, tunnel, VPS, and model access problems.

See Rescue Session

Read next

Strix Halo vs Mac Studio M4 Max 128GB for Local LLMs: Which Unified Memory Box?
Compare AMD Strix Halo (Ryzen AI Max+ 395) and Mac Studio M4 Max 128GB for local LLMs: 256 vs 546 GB/s bandwidth, decode speed, ROCm vs MLX, August 2026 prices.
Which Ryzen AI Max+ 395 Mini PC Should You Buy? (August 2026)
GMKtec EVO-X2 vs Framework Desktop vs Beelink GTR9 Pro vs Minisforum MS-S1 MAX. Same APU, same ~256 GB/s, $1,959 to $4,349 — and the cheapest ones are out of stock. What to actually buy in August 2026.
Jetson Thor vs DGX Spark (August 2026): Which NVIDIA 128GB Box Is For You?
Jetson AGX Thor and DGX Spark both carry 128GB of LPDDR5X at exactly 273 GB/s, so they generate tokens at the same ceiling. Thor lists at $3,499 against Spark's $4,699. The real decision is deploy versus develop, not TOPS — and street pricing reverses the MSRP gap.
MacBook Pro M4 Max for AI: 36GB vs 128GB (Which RAM for Local LLMs?)
36GB or 128GB M4 Max for local AI? The 36GB config ships on the 14-core M4 Max at 410 GB/s; 128GB requires the 16-core chip at 546 GB/s. 36GB runs Qwen 3.6 27B Q8 and Laguna XS 2.1; 128GB is the only way to run gpt-oss 120B or Llama 4 Scout locally.