← All guides

Can I Run LTX-2.5 Locally? VRAM by GPU (2026)

Lightricks gives two different answers. Its model page says 'Min VRAM: 16GB'. Its own system requirements page says a minimum of 32GB. Both are true for a different setup. We read every LTX-2.5 file size from the Hugging Face API on 24 September 2026, checked ComfyUI, LTX Desktop, Ollama and the community GGUF repos, and found why the numbers disagree: the 12B text encoder, and where it runs.

Bottom Line

  • Yes, you can run LTX-2.5 locally. A 24GB card holds the official INT8 transformer. A 16GB card works with NVFP4 or a community GGUF, plus weight streaming to system RAM.
  • Lightricks gives two minimums. The LTX-2.5 model page says “Min VRAM: 16GB”. The system requirements page says “minimum 32GB+ VRAM” and recommends an 80GB A100 or H100.
  • The text encoder explains the gap. LTX-2.5 needs a custom Gemma 4 12B encoder. It is 26.26 GB at BF16 and 15.37 GB at INT8, about 40% of every official file set.
  • LTX Desktop’s 16GB path encodes your prompt in the cloud by default. Its README calls cloud text encoding free and “highly recommended to speed up inference and save memory”. The local encoder is an extra download.
  • Ollama and LM Studio do not run it. Use ComfyUI (native since v0.32.0), LTX Desktop, or the Lightricks Python pipelines.
  • License: free commercial use under $10M annual revenue.

Minimum VRAM by format, from the file sizes below (our estimate, see the method):

Format (transformer + encoder + 2 VAEs)Files totalLargest single phaseMinimum card
BF16 + BF16 encoder (official)70.12 GB43.86 GB (sampling)48GB+, or 32GB with offload
INT8 + INT8 encoder (official, ComfyUI)38.71 GB23.34 GB (sampling)24GB
NVFP4 + INT8 encoder (official)35.93 GB20.56 GB (sampling)24GB, or 16GB with streaming
Q4_K_S GGUF + Q4_K_M encoder (community)24.10 GB15.69 GB (sampling)16GB, tight
Q3_K_M GGUF + Q2_K encoder (community)19.33 GB13.37 GB (sampling)16GB
Q2_K GGUF + Q2_K encoder (community)16.63 GB10.67 GB (sampling)12GB, low quality

What the Model Is

Figures come from the Lightricks/LTX-2.5 model card, the LTX-2 GitHub README and the Hugging Face API, read on 24 September 2026.

SpecValue
Transformer22B parameters, dev and distilled versions
Text encoderGemma 4 12B, fine-tuned by Lightricks, with a text projection bundled in
Video decodersDiffusion decoder (“higher quality, heavier”) or conv VAE (“faster, lighter”)
AudioGenerated in the same pass, with its own audio VAE and vocoder
Distilled stepsFixed 8-step schedule; the dev model needs about 30 steps (per a reply in HF discussion #35)
Frame countMust be 8n+1 (1, 9, 17 … 121). 121 frames is about 5 s at 24 fps
ResolutionWidth and height divisible by 32. 4K is 3840x2176 through the x2 spatial upscaler
LicenseLTX-2.x Community License, free under $10M annual revenue
Released11 August 2026 (the HF repo was created 23 July, gated)

The encoder is not stock Gemma. A Lightricks staff member wrote that it “can’t be replaced by a standard gemma” (HF discussion #43). The GitHub README says the loader checks the encoder version. So the popular Gemma 4 12B GGUFs for chat will not work here.

Every File Size That Matters

Sizes come from the Hugging Face API with ?blobs=true, read on 24 September 2026. GB means 10^9 bytes.

Transformer

FileSourceSize (GB)
Distilled BF16Lightricks42.02
Dev BF16Lightricks42.02
Distilled INT8 ConvRot (ComfyUI)Lightricks21.50
Dev INT8 ConvRot (ComfyUI)Lightricks21.50
Distilled NVFP4Lightricks18.72
Q8_0 GGUFAbiray, vantagewithai23.60
Q6_K GGUFAbiray, vantagewithai18.62
Q5_K_M GGUFAbiray, vantagewithai18.12
Q4_K_M GGUFAbiray, vantagewithai15.69
Q4_K_M / Q4_K_S GGUFrealrebelai15.09 / 13.85
Q3_K_M GGUFAbiray / realrebelai12.92 / 11.53
Q2_K GGUFvantagewithai / realrebelai12.13 / 8.83

The GGUFs are large for a 22B model. vantagewithai says its files use “mixed-precision levels” and are larger as a result. realrebelai keeps the audio-video gate layers at high precision because quantizing them “desyncs or degrades” the audio. Same quant name, different sizes: check the repo before you download.

Text encoder (Gemma 4 12B, LTX version)

FileSourceSize (GB)
BF16Lightricks26.26
INT8 ConvRot (ComfyUI)Lightricks15.37
Q5_K_M GGUFelix3r (gated)9.51
Q4_K_M GGUFelix3r8.41
Q2_K GGUFelix3r5.96

VAEs and extras

FileSize (GB)
Video VAE (diffusion decoder)1.47
Video VAE (conv)1.45
Audio VAE0.37
Spatial upscaler x2 (optional)1.00
Temporal upscaler x2 (optional)0.26
Distilled LoRA (to speed up dev)8.90

The fact the “16GB minimum” leaves out

The text encoder is 37-43% of every official file set: 26.26 of 70.12 GB at BF16, 15.37 of 38.71 GB at INT8, 15.37 of 35.93 GB with NVFP4. The INT8 encoder alone is 15.37 GB. That is almost a full 16GB card before the video model loads.

The LTX Desktop README shows how Lightricks gets to 16GB. The app runs LTX 2.5 Fast locally on NVIDIA cards with “≥16GB VRAM”. Below that, it switches to paid API mode. For text encoding, the README recommends a free LTX API key, which sends your prompts to the LTX API. A “Local Text Encoder” is an extra download in settings. So the default 16GB experience is local video, cloud prompt encoding. That is a fair design, but it is not the same as fully offline.

Fully offline on 16GB means you pick a smaller encoder. The Q4_K_M GGUF encoder is 8.41 GB, about a third of the BF16 file.

Decoder Table: Your GPU, Your Files

This table is our estimate. Method: ComfyUI encodes the prompt first, frees the encoder, then loads the transformer for sampling. The Smeltcore RTX 4090 recipe describes this order. So the peak is the larger phase, not the sum of all files. The sampling phase is the transformer plus both VAEs. Activations add more on top, and they grow with resolution and clip length.

VRAMTransformerText encoderExpected experience
8GBNone fits residentQ2_K GGUF (5.96 GB)Not a practical target. LTX Desktop drops to API mode below 16GB. We found no confirmed 8GB run.
12GBQ2_K GGUF (8.83 GB)Q2_K or Q4_K_M GGUF, offloaded to CPUFits on paper at 10.67 GB for sampling. realrebelai says Q2_K quality “drops sharply”. We could not confirm a 12GB report.
16GBQ3_K_M (11.53-12.92 GB) or Q4_K_S (13.85 GB) GGUF; NVFP4 with streamingQ4_K_M GGUF (8.41 GB), or INT8 (15.37 GB) with offloadWorks. Keep clips short. Plan on 32GB+ of system RAM for offloaded weights.
24GBINT8 ConvRot (21.50 GB)INT8 ConvRot (15.37 GB)Official ComfyUI files, transformer resident. Sampling phase is 23.34 GB of files, so headroom is small. Use the conv VAE and short clips.
32GBINT8 (21.50 GB) resident, or BF16 (42.02 GB) with offloadINT8 or BF16Room for longer clips at 720p. The official Python path lists 32GB as its minimum.

NVFP4 on older cards. The NVFP4 file is the smallest official transformer, but FP4 math needs a Blackwell GPU (RTX 50 series). The ComfyUI NVFP4 post says the model still runs on older cards, but sampling “may actually be up to 2x slower than fp8”. On an RTX 3090, use INT8.

Measured Reports, With Sources

We did not run these tests. Each is one machine and one setup.

GPUSetupResultSource
RTX 5090 32GBComfyUI official workflow, distilled BF16 with offloading720p, 8 s image-to-video in about 80 sTrueNorthAI (note.com), 12 Aug 2026
RTX 5090 32GBDistilled INT8 ConvRotNormal output up to about 10 s at 720p. A 15 s 720p text-to-video run ran out of memorysame
RTX 5060 Ti 16GB, 32 GB RAMComfyUI v0.32.0, distilled NVFP4, INT8 encoder, two-stage DFR1344x768, 24 fps, about 4.4 s clip in about 170 s. System RAM use above 92%Reddit r/comfyui post, relayed by AI Creative Log (note.com), 15 Aug 2026. We could not open the Reddit original
MacBook Pro M4 Max, 36 GBDefault ComfyUI workflow”Worked just fine.” No timing givenHF discussion #58, 22 Aug 2026

Two details from these reports matter:

  1. Clip length, not resolution, breaks 32GB first. The RTX 5090 handled about 10 s at 720p with INT8, then failed at 15 s. Start at 5 s (121 frames) and grow from there.
  2. System RAM is part of the budget. The 16GB report used over 92% of 32 GB of RAM. Lightricks lists 32 GB of RAM as the minimum and 64 GB as recommended.

The same TrueNorthAI test hit a “dimension mismatch” error with the NVFP4 file on 12 August. The 5060 Ti report three days later ran NVFP4. Update ComfyUI before you test it.

Which Tools Run It

We checked each runtime on 24 September 2026.

RuntimeStatus
ComfyUINative. “Add support for LTX 2.5” (PR #15499) shipped in v0.32.0 on 11 Aug 2026. Official templates use the distilled model.
ComfyUI + GGUFWorks with the stock ComfyUI-GGUF Unet loader on realrebelai’s files. That repo explains that a plain GGUF conversion drops the model config and fails with shape errors.
LTX Desktop (beta)Yes, LTX 2.5 Fast. Windows or Linux with NVIDIA and 16GB+ VRAM, or Apple Silicon with 15GB+ free RAM. Needs 160GB+ free disk. LTX 2.5 Pro is API only.
Lightricks Python pipelinesYes. The quick start download is “roughly 66 GiB” of BF16 files. --quantization fp8-cast --offload cpu is the documented low-memory option.
Apple SiliconLTX Desktop runs on MPS. MLX ports exist (for example mlx-community/ltx-2.5-mlx). A ComfyUI MPS bug that produced black videos was reported as issue #15804.
OllamaNo. See below.
LM Studio, llama.cppNo. These run text and vision LLMs. LM Studio lists no LTX model.

Does Ollama run LTX-2.5?

No. On 24 September 2026, ollama.com/library/ltx-2.5 and ollama.com/x/ltx-2.5 both return 404. An Ollama search for “ltx” returns no LTX model. The confusion comes from the file name: the community files are GGUF, and Ollama runs GGUF. But these GGUFs hold a video diffusion transformer that only ComfyUI’s GGUF loader understands.

Some ComfyUI workflows call Ollama to rewrite prompts before LTX-2.5 runs. In that setup, Ollama writes text. It does not make the video.

Which GPU to Buy for It

Prices come from our tracked price file, as of September 2026 unless marked.

  • Best fit: 24GB. A used RTX 3090 24GB (about $1,453 average used) is the cheapest card that holds the official 21.50 GB INT8 transformer without streaming. No FP4 support, so skip the NVFP4 file on it.
  • Longer clips: 32GB. The RTX 5090 32GB (from $4,299) meets the official 32GB minimum and ran about 10 s at 720p in the report above. It also has native FP4 for the NVFP4 file.
  • Budget: 16GB. The RTX 5060 Ti 16GB ($589-805 as of August 2026) is Blackwell, so NVFP4 runs at full speed. Expect minutes per clip, short clips, and heavy system RAM use. Pair it with 32GB of RAM or more.

AMD, Intel and Mac: LTX Desktop does not list AMD or Intel GPUs for local generation, and we found no measured AMD report. Macs run it through LTX Desktop or MLX, but we found no timed Mac benchmark. For other tiers, see the tested hardware list.

Honest Caveats

  1. We did not run the model. Every speed number is a community report with a named source.
  2. The decoder table is arithmetic. It uses file sizes and the encode-then-sample order ComfyUI uses. Activations for long or high-resolution clips add several GB we cannot pin down.
  3. Lightricks’ two minimums describe different setups. 16GB is LTX Desktop with its optimizations and, by default, cloud text encoding. 32GB is the official Python pipeline.
  4. Community GGUFs vary. The same quant name differs by up to 3.3 GB between repos (Q2_K: 8.83 vs 12.13 GB). The realrebelai card lists approximate sizes that are smaller than the API values; we used the API.
  5. The weights are gated. You must accept the terms on Hugging Face before download. The GGUF encoder repo is gated too.
  6. The license has a revenue cap. Free under $10M annual revenue. Read the LTX-2.x Community License before commercial use.

Sources

See Also

Need OpenClaw fixed live?

Remote rescue sessions for gateway, auth, tunnel, VPS, and model access problems.

See Rescue Session

Read next

Is a Used RTX 3090 Worth $1,450 for Local LLMs? (2026)
A used RTX 3090 sells for a $1,463 average as of September 2026, about 97% of its $1,499 launch price in 2020. It is no longer the cheapest VRAM per GB, but it is still the cheapest memory bandwidth. A verdict per buyer, what $1,450 buys instead, and the rent-vs-buy hours.
Can I Run Qwen-Image 2.1 Locally? Yes, on 12GB
Qwen-Image 2.1 is a 7B image model, but its Qwen3-VL 8B text encoder is bigger than the image model. A Q4 GGUF stack is 9.91 GB, so 12GB works. 16GB runs INT8 at about 20 s per image. Ollama does not run it.
Can I Run MiniMax H3 Locally? Yes, 16GB Works (2026)
MiniMax H3 is a 33B audio-video model plus a Qwen3-VL-32B text encoder. A 16GB card runs ComfyUI's pruned INT8 files with offload. 24GB holds the video model. System RAM and the license territory matter more than most guides say.
Best Qwen Model for 16GB VRAM (2026): Qwen3.8 27B Wins
Best Qwen model for 16GB VRAM: Qwen3.8-27B at 3.5 bits (10.96 GiB). The Q4_K_M 27B files do not fit. Byte-exact GGUF sizes, KV math, and the best Qwen for 8, 12, 24 and 32GB.