Jetson Thor vs DGX Spark (August 2026): Which NVIDIA 128GB Box Is For You?
Jetson AGX Thor and DGX Spark get confused constantly, because the spec sheets rhyme: both are small NVIDIA boxes, both say Blackwell, both carry 128GB of unified LPDDR5X. The number that decides local LLM speed is identical on both — 273 GB/s on a 256-bit bus — which means they generate tokens at the same ceiling no matter what the TOPS figures suggest. What actually separates them is the software image and what NVIDIA built each one to do. Thor is a deployment target for robots. Spark is a development machine. Most people asking this question want the one they are not looking at.
Sizing a 128GB box for OpenClaw?
See our AI training options. We'll work out whether you need this tier at all before you spend $4,000.
Bottom Line (August 2026)
- They generate tokens at the same speed ceiling. Both carry 128GB of 256-bit LPDDR5X at 273 GB/s. Decode is bandwidth-bound. The TOPS gap does not change that.
- Thor is for deploying. Spark is for developing. NVIDIA built Thor as a physical-AI platform for robots that run finished models continuously at 40-130W. Spark is a desktop machine for building them.
- MSRP favours Thor by $1,200 — $3,499 against $4,699 — but street pricing reverses it. Thor kits have listed at $5,499.
- Buy Thor only if you are building an embedded system. For a desktop LLM box, Spark or the ASUS Ascent GX10 fit better, and a Ryzen AI Max+ 395 box is far cheaper.
- Ignore the TOPS comparison. Thor’s 2070 TFLOPS is a sparse FP4 figure. Sparse numbers run roughly double their dense equivalent, and neither predicts tokens per second.
The Specs Side by Side
| Jetson AGX Thor (T5000) | DGX Spark (GB10) | |
|---|---|---|
| Memory | 128GB LPDDR5X, 256-bit | 128GB LPDDR5X, 256-bit |
| Memory bandwidth | 273 GB/s | 273 GB/s |
| GPU | 2,560-core Blackwell, 96 5th-gen Tensor Cores | 6,144 CUDA cores, Blackwell |
| Quoted AI compute | 2,070 TFLOPS (sparse FP4) | ~1 petaflop |
| CPU | 14-core Arm Neoverse-V3AE | 20-core Arm (10x Cortex-X925, 10x A725) |
| Power | 40W - 130W | Desktop appliance, ~1.2 kg |
| List price | $3,499 | $4,699 (raised from $3,999) |
| Seen at | $5,499 (Aug 2026 listings) | Partner boxes often slightly less |
| Built for | Physical AI, robotics, edge deployment | Local model development |
Specs from NVIDIA’s published Jetson Thor module table and DGX Spark documentation. Prices as of August 2026 and volatile.
The row that matters for local LLMs is the second one, and it is a tie.
Why the Bandwidth Tie Settles the Speed Question
When a language model writes you an answer, it produces one token at a time, and producing each token requires reading the active weights out of memory. Arithmetic is cheap by comparison. This is why memory bandwidth — not TOPS, not core count — is the number that predicts tokens per second on a machine like this.
Both boxes read memory at 273 GB/s. So for the same model at the same quant, expect the same decode speed, within the noise of whatever runtime each is using.
Where Thor’s extra compute genuinely helps:
- Prefill — chewing through a long prompt before the first token appears. This is compute-heavy and parallel, so more TFLOPS shows up as a faster time-to-first-token.
- Vision and multimodal — image encoders are compute-bound in a way text decode is not.
- Batching — serving several requests at once shifts the balance toward compute.
- Anything real-time and sensor-driven, which is exactly what Thor was designed for.
And a caution on the numbers themselves: Thor’s 2,070 TFLOPS is explicitly a sparse FP4 figure. Sparse figures assume structured sparsity in the weights and land at roughly twice the dense equivalent. Spark’s “1 petaflop” is quoted at FP4 as well. Comparing headline AI-compute numbers across vendors — or even across two NVIDIA product lines — is only meaningful when both are stated the same way, and marketing pages rarely make that easy. Treat the compute row as “Thor has more, by some amount less than the ratio suggests,” and treat the bandwidth row as exact.
The Real Split: Deploy vs Develop
NVIDIA is unusually clear about this, and most comparison pages ignore it in favour of the spec table.
Jetson Thor is the platform for physical AI. NVIDIA describes it as a supercomputer for generative reasoning and multimodal, multisensor processing, aimed at building next-generation humanoid robots — object manipulation, navigation, following complex instructions. The 40W-130W configurable power envelope, the 14-core embedded-class Neoverse-V3AE CPU, and the developer kit’s 4x 25GbE networking all point the same direction: a machine that goes inside something and runs a finished model continuously against live sensors.
DGX Spark is a development machine. It sits on a desk, runs the NVIDIA AI software stack, and exists so you can build and test against the same CUDA environment your datacenter deployment will use. Its honest case has never been raw speed — it is CUDA parity with your deployment target.
That framing answers the question for almost everyone who asks it:
- Building a robot, drone, or embedded system that must run models in the field? Thor. It is the only one of the two designed for that, and nothing about Spark replaces it.
- Want a quiet 128GB box on your desk to run and develop local models? Spark, or a partner GB10 box. Thor will do it, but you will be fighting a robotics-oriented image to get there and paying street prices that erase the MSRP saving.
- Just want to run big models locally as cheaply as possible? Neither. See the alternatives below.
The Software Point People Miss
On the Spark, the stack is worth more than the silicon. The same box benchmarks 11.7 tok/s on Ollama, 38.5 on tuned llama.cpp, and around 50 on SGLang for gpt-oss 120B — a 4x spread on identical hardware.
Carry that finding into this comparison and it reframes everything. If the runtime is worth 4x on the same box, then a hardware difference of a few hundred GB/s of nominal compute — on two machines with identical memory bandwidth — is not the variable you should be optimising. The box you can most easily get a well-tuned runtime onto will beat the box with the better spec sheet. For LLM work specifically, that is Spark, because that is the workload its image was built around.
This is also the honest argument against buying a Thor as a bargain Spark. On paper you save $1,200 and gain compute. In practice you are choosing the machine with less LLM-tuned software support, for a workload where software support has been measured at 4x.
What to Buy
If you want the NVIDIA desktop box for local model work, the NVIDIA DGX Spark 128GB GB10 Grace Blackwell is the reference option at $4,699.
The ASUS Ascent GX10 is the same GB10 silicon and the same 128GB at a lower price — the better value of the two if you do not specifically need NVIDIA's own box and its DGX OS image.
If you only want 128GB of unified memory to run models and do not need CUDA, a Ryzen AI Max+ 395 (Strix Halo) 128GB box holds the same class of model for roughly half the money, at lower bandwidth and with ROCm rather than CUDA.
Jetson AGX Thor is not in our affiliate catalog, so this is a plain recommendation with nothing attached: if you are genuinely building physical AI, buy the developer kit at the closest price you can find to its $3,499 list, and treat any $5,000+ listing as a reason to wait.
What Nobody Else Will Tell You
Every comparison of these two boxes leads with the compute figures, because 2,070 against 1,000 is a satisfying headline. It is also close to irrelevant for the workload most readers have in mind, and the number that is relevant is a dead tie that no one prints.
The deeper point is that NVIDIA did not build two competing boxes here. It built one machine to develop models and one machine to deploy them, gave them the same memory subsystem, and let the software images diverge. People shopping on the spec table see a $1,200 discount and 2x the TOPS and reasonably conclude Thor is the better buy. They are comparing the wrong column, on prices that have already inverted, for a decision that memory bandwidth settled before either number mattered.
See Also
- Is the DGX Spark Worth It? — the full verdict at $4,699, and the 4x runtime spread on identical hardware
- Best Models for DGX Spark — what actually runs well on GB10, by cluster size
- Best Models for the ASUS Ascent GX10 — same silicon, lower price, what changes
- DGX Spark vs Strix Halo — two 128GB boxes $2,100 apart
- Cheapest Way to Run a 70B Model Locally — why “fits” and “runs well” are different questions at this tier
Need OpenClaw fixed live?
Remote rescue sessions for gateway, auth, tunnel, VPS, and model access problems.
See Rescue Session