Can I Run LTX-2.5 Locally? VRAM by GPU (2026)
Lightricks gives two different answers. Its model page says 'Min VRAM: 16GB'. Its own system requirements page says a minimum of 32GB. Both are true for a different setup. We read every LTX-2.5 file size from the Hugging Face API on 24 September 2026, checked ComfyUI, LTX Desktop, Ollama and the community GGUF repos, and found why the numbers disagree: the 12B text encoder, and where it runs.
Bottom Line
- Yes, you can run LTX-2.5 locally. A 24GB card holds the official INT8 transformer. A 16GB card works with NVFP4 or a community GGUF, plus weight streaming to system RAM.
- Lightricks gives two minimums. The LTX-2.5 model page says “Min VRAM: 16GB”. The system requirements page says “minimum 32GB+ VRAM” and recommends an 80GB A100 or H100.
- The text encoder explains the gap. LTX-2.5 needs a custom Gemma 4 12B encoder. It is 26.26 GB at BF16 and 15.37 GB at INT8, about 40% of every official file set.
- LTX Desktop’s 16GB path encodes your prompt in the cloud by default. Its README calls cloud text encoding free and “highly recommended to speed up inference and save memory”. The local encoder is an extra download.
- Ollama and LM Studio do not run it. Use ComfyUI (native since v0.32.0), LTX Desktop, or the Lightricks Python pipelines.
- License: free commercial use under $10M annual revenue.
Minimum VRAM by format, from the file sizes below (our estimate, see the method):
| Format (transformer + encoder + 2 VAEs) | Files total | Largest single phase | Minimum card |
|---|---|---|---|
| BF16 + BF16 encoder (official) | 70.12 GB | 43.86 GB (sampling) | 48GB+, or 32GB with offload |
| INT8 + INT8 encoder (official, ComfyUI) | 38.71 GB | 23.34 GB (sampling) | 24GB |
| NVFP4 + INT8 encoder (official) | 35.93 GB | 20.56 GB (sampling) | 24GB, or 16GB with streaming |
| Q4_K_S GGUF + Q4_K_M encoder (community) | 24.10 GB | 15.69 GB (sampling) | 16GB, tight |
| Q3_K_M GGUF + Q2_K encoder (community) | 19.33 GB | 13.37 GB (sampling) | 16GB |
| Q2_K GGUF + Q2_K encoder (community) | 16.63 GB | 10.67 GB (sampling) | 12GB, low quality |
What the Model Is
Figures come from the Lightricks/LTX-2.5 model card, the LTX-2 GitHub README and the Hugging Face API, read on 24 September 2026.
| Spec | Value |
|---|---|
| Transformer | 22B parameters, dev and distilled versions |
| Text encoder | Gemma 4 12B, fine-tuned by Lightricks, with a text projection bundled in |
| Video decoders | Diffusion decoder (“higher quality, heavier”) or conv VAE (“faster, lighter”) |
| Audio | Generated in the same pass, with its own audio VAE and vocoder |
| Distilled steps | Fixed 8-step schedule; the dev model needs about 30 steps (per a reply in HF discussion #35) |
| Frame count | Must be 8n+1 (1, 9, 17 … 121). 121 frames is about 5 s at 24 fps |
| Resolution | Width and height divisible by 32. 4K is 3840x2176 through the x2 spatial upscaler |
| License | LTX-2.x Community License, free under $10M annual revenue |
| Released | 11 August 2026 (the HF repo was created 23 July, gated) |
The encoder is not stock Gemma. A Lightricks staff member wrote that it “can’t be replaced by a standard gemma” (HF discussion #43). The GitHub README says the loader checks the encoder version. So the popular Gemma 4 12B GGUFs for chat will not work here.
Every File Size That Matters
Sizes come from the Hugging Face API with ?blobs=true, read on 24 September 2026. GB means 10^9 bytes.
Transformer
| File | Source | Size (GB) |
|---|---|---|
| Distilled BF16 | Lightricks | 42.02 |
| Dev BF16 | Lightricks | 42.02 |
| Distilled INT8 ConvRot (ComfyUI) | Lightricks | 21.50 |
| Dev INT8 ConvRot (ComfyUI) | Lightricks | 21.50 |
| Distilled NVFP4 | Lightricks | 18.72 |
| Q8_0 GGUF | Abiray, vantagewithai | 23.60 |
| Q6_K GGUF | Abiray, vantagewithai | 18.62 |
| Q5_K_M GGUF | Abiray, vantagewithai | 18.12 |
| Q4_K_M GGUF | Abiray, vantagewithai | 15.69 |
| Q4_K_M / Q4_K_S GGUF | realrebelai | 15.09 / 13.85 |
| Q3_K_M GGUF | Abiray / realrebelai | 12.92 / 11.53 |
| Q2_K GGUF | vantagewithai / realrebelai | 12.13 / 8.83 |
The GGUFs are large for a 22B model. vantagewithai says its files use “mixed-precision levels” and are larger as a result. realrebelai keeps the audio-video gate layers at high precision because quantizing them “desyncs or degrades” the audio. Same quant name, different sizes: check the repo before you download.
Text encoder (Gemma 4 12B, LTX version)
| File | Source | Size (GB) |
|---|---|---|
| BF16 | Lightricks | 26.26 |
| INT8 ConvRot (ComfyUI) | Lightricks | 15.37 |
| Q5_K_M GGUF | elix3r (gated) | 9.51 |
| Q4_K_M GGUF | elix3r | 8.41 |
| Q2_K GGUF | elix3r | 5.96 |
VAEs and extras
| File | Size (GB) |
|---|---|
| Video VAE (diffusion decoder) | 1.47 |
| Video VAE (conv) | 1.45 |
| Audio VAE | 0.37 |
| Spatial upscaler x2 (optional) | 1.00 |
| Temporal upscaler x2 (optional) | 0.26 |
| Distilled LoRA (to speed up dev) | 8.90 |
The fact the “16GB minimum” leaves out
The text encoder is 37-43% of every official file set: 26.26 of 70.12 GB at BF16, 15.37 of 38.71 GB at INT8, 15.37 of 35.93 GB with NVFP4. The INT8 encoder alone is 15.37 GB. That is almost a full 16GB card before the video model loads.
The LTX Desktop README shows how Lightricks gets to 16GB. The app runs LTX 2.5 Fast locally on NVIDIA cards with “≥16GB VRAM”. Below that, it switches to paid API mode. For text encoding, the README recommends a free LTX API key, which sends your prompts to the LTX API. A “Local Text Encoder” is an extra download in settings. So the default 16GB experience is local video, cloud prompt encoding. That is a fair design, but it is not the same as fully offline.
Fully offline on 16GB means you pick a smaller encoder. The Q4_K_M GGUF encoder is 8.41 GB, about a third of the BF16 file.
Decoder Table: Your GPU, Your Files
This table is our estimate. Method: ComfyUI encodes the prompt first, frees the encoder, then loads the transformer for sampling. The Smeltcore RTX 4090 recipe describes this order. So the peak is the larger phase, not the sum of all files. The sampling phase is the transformer plus both VAEs. Activations add more on top, and they grow with resolution and clip length.
| VRAM | Transformer | Text encoder | Expected experience |
|---|---|---|---|
| 8GB | None fits resident | Q2_K GGUF (5.96 GB) | Not a practical target. LTX Desktop drops to API mode below 16GB. We found no confirmed 8GB run. |
| 12GB | Q2_K GGUF (8.83 GB) | Q2_K or Q4_K_M GGUF, offloaded to CPU | Fits on paper at 10.67 GB for sampling. realrebelai says Q2_K quality “drops sharply”. We could not confirm a 12GB report. |
| 16GB | Q3_K_M (11.53-12.92 GB) or Q4_K_S (13.85 GB) GGUF; NVFP4 with streaming | Q4_K_M GGUF (8.41 GB), or INT8 (15.37 GB) with offload | Works. Keep clips short. Plan on 32GB+ of system RAM for offloaded weights. |
| 24GB | INT8 ConvRot (21.50 GB) | INT8 ConvRot (15.37 GB) | Official ComfyUI files, transformer resident. Sampling phase is 23.34 GB of files, so headroom is small. Use the conv VAE and short clips. |
| 32GB | INT8 (21.50 GB) resident, or BF16 (42.02 GB) with offload | INT8 or BF16 | Room for longer clips at 720p. The official Python path lists 32GB as its minimum. |
NVFP4 on older cards. The NVFP4 file is the smallest official transformer, but FP4 math needs a Blackwell GPU (RTX 50 series). The ComfyUI NVFP4 post says the model still runs on older cards, but sampling “may actually be up to 2x slower than fp8”. On an RTX 3090, use INT8.
Measured Reports, With Sources
We did not run these tests. Each is one machine and one setup.
| GPU | Setup | Result | Source |
|---|---|---|---|
| RTX 5090 32GB | ComfyUI official workflow, distilled BF16 with offloading | 720p, 8 s image-to-video in about 80 s | TrueNorthAI (note.com), 12 Aug 2026 |
| RTX 5090 32GB | Distilled INT8 ConvRot | Normal output up to about 10 s at 720p. A 15 s 720p text-to-video run ran out of memory | same |
| RTX 5060 Ti 16GB, 32 GB RAM | ComfyUI v0.32.0, distilled NVFP4, INT8 encoder, two-stage DFR | 1344x768, 24 fps, about 4.4 s clip in about 170 s. System RAM use above 92% | Reddit r/comfyui post, relayed by AI Creative Log (note.com), 15 Aug 2026. We could not open the Reddit original |
| MacBook Pro M4 Max, 36 GB | Default ComfyUI workflow | ”Worked just fine.” No timing given | HF discussion #58, 22 Aug 2026 |
Two details from these reports matter:
- Clip length, not resolution, breaks 32GB first. The RTX 5090 handled about 10 s at 720p with INT8, then failed at 15 s. Start at 5 s (121 frames) and grow from there.
- System RAM is part of the budget. The 16GB report used over 92% of 32 GB of RAM. Lightricks lists 32 GB of RAM as the minimum and 64 GB as recommended.
The same TrueNorthAI test hit a “dimension mismatch” error with the NVFP4 file on 12 August. The 5060 Ti report three days later ran NVFP4. Update ComfyUI before you test it.
Which Tools Run It
We checked each runtime on 24 September 2026.
| Runtime | Status |
|---|---|
| ComfyUI | Native. “Add support for LTX 2.5” (PR #15499) shipped in v0.32.0 on 11 Aug 2026. Official templates use the distilled model. |
| ComfyUI + GGUF | Works with the stock ComfyUI-GGUF Unet loader on realrebelai’s files. That repo explains that a plain GGUF conversion drops the model config and fails with shape errors. |
| LTX Desktop (beta) | Yes, LTX 2.5 Fast. Windows or Linux with NVIDIA and 16GB+ VRAM, or Apple Silicon with 15GB+ free RAM. Needs 160GB+ free disk. LTX 2.5 Pro is API only. |
| Lightricks Python pipelines | Yes. The quick start download is “roughly 66 GiB” of BF16 files. --quantization fp8-cast --offload cpu is the documented low-memory option. |
| Apple Silicon | LTX Desktop runs on MPS. MLX ports exist (for example mlx-community/ltx-2.5-mlx). A ComfyUI MPS bug that produced black videos was reported as issue #15804. |
| Ollama | No. See below. |
| LM Studio, llama.cpp | No. These run text and vision LLMs. LM Studio lists no LTX model. |
Does Ollama run LTX-2.5?
No. On 24 September 2026, ollama.com/library/ltx-2.5 and ollama.com/x/ltx-2.5 both return 404. An Ollama search for “ltx” returns no LTX model. The confusion comes from the file name: the community files are GGUF, and Ollama runs GGUF. But these GGUFs hold a video diffusion transformer that only ComfyUI’s GGUF loader understands.
Some ComfyUI workflows call Ollama to rewrite prompts before LTX-2.5 runs. In that setup, Ollama writes text. It does not make the video.
Which GPU to Buy for It
Prices come from our tracked price file, as of September 2026 unless marked.
- Best fit: 24GB. A used RTX 3090 24GB (about $1,453 average used) is the cheapest card that holds the official 21.50 GB INT8 transformer without streaming. No FP4 support, so skip the NVFP4 file on it.
- Longer clips: 32GB. The RTX 5090 32GB (from $4,299) meets the official 32GB minimum and ran about 10 s at 720p in the report above. It also has native FP4 for the NVFP4 file.
- Budget: 16GB. The RTX 5060 Ti 16GB ($589-805 as of August 2026) is Blackwell, so NVFP4 runs at full speed. Expect minutes per clip, short clips, and heavy system RAM use. Pair it with 32GB of RAM or more.
AMD, Intel and Mac: LTX Desktop does not list AMD or Intel GPUs for local generation, and we found no measured AMD report. Macs run it through LTX Desktop or MLX, but we found no timed Mac benchmark. For other tiers, see the tested hardware list.
Honest Caveats
- We did not run the model. Every speed number is a community report with a named source.
- The decoder table is arithmetic. It uses file sizes and the encode-then-sample order ComfyUI uses. Activations for long or high-resolution clips add several GB we cannot pin down.
- Lightricks’ two minimums describe different setups. 16GB is LTX Desktop with its optimizations and, by default, cloud text encoding. 32GB is the official Python pipeline.
- Community GGUFs vary. The same quant name differs by up to 3.3 GB between repos (Q2_K: 8.83 vs 12.13 GB). The realrebelai card lists approximate sizes that are smaller than the API values; we used the API.
- The weights are gated. You must accept the terms on Hugging Face before download. The GGUF encoder repo is gated too.
- The license has a revenue cap. Free under $10M annual revenue. Read the LTX-2.x Community License before commercial use.
Sources
- Lightricks/LTX-2.5 model card: 22B transformer, Gemma 4 12B encoder, frame rules, 4K upscaling, 8-step distilled schedule, license tiers
- Hugging Face API: Lightricks/LTX-2.5: every official file size, repo creation date
- LTX system requirements: 32GB+ VRAM minimum, A100/H100 recommended, 32/64 GB RAM
- LTX-2.5 model page: “Min VRAM: 16GB”
- LTX-2 GitHub README: 66 GiB quick start, fp8-cast and offload flags, encoder version check
- LTX Desktop README: 16GB VRAM local mode, 15GB free RAM on Mac, cloud text encoding
- Abiray/LTX-2.5-Distilled-GGUF, vantagewithai/LTX-2.5-GGUF, realrebelai/LTX-2.5_GGUFs: GGUF sizes, loader notes
- elix3r/gemma4-12b-with-proj-ltx-2.5-GGUF: encoder GGUF sizes
- ComfyUI v0.32.0 release: LTX 2.5 support, 11 Aug 2026
- ComfyUI NVFP4 post: FP4 needs Blackwell, up to 2x slower elsewhere
- Smeltcore RTX 4090 recipe: encode-then-sample loading order
- TrueNorthAI RTX 5090 test: 5090 timings and the 15 s OOM
- AI Creative Log roundup: RTX 5060 Ti 16GB report
- HF discussions #35, #43, #58: dev steps, custom encoder, Mac run
- Datanorth release report: 11 August 2026 release
See Also
- Can I run Qwen-Image 2.1 locally?: the image-model sibling, where the text encoder is also the bigger half
- Best local LLM for an RTX 3090: what else the 24GB card runs
- Best local LLM for an RTX 5090: the 32GB tier
- Best local LLM for an RTX 5060 Ti 16GB: the Blackwell 16GB budget card
- Is 16GB of VRAM still enough in 2026?: the tier-level verdict
- Is 32GB of VRAM enough in 2026?: what the step to 32GB buys
- GGUF quant names explained: what Q4_K_M, Q3_K_M and Q8_0 mean
- Which used RTX 3090 to buy: how to pick the 24GB card
- Can I run MiniMax H3 locally?: the 33B audio-video model, where the Qwen3-VL-32B encoder outweighs the pruned video model
Need OpenClaw fixed live?
Remote rescue sessions for gateway, auth, tunnel, VPS, and model access problems.
See Rescue Session