Local LLM Electricity Cost vs API (July 2026): The Break-Even Math Including Power and Depreciation
Nobody puts the power bill in the spreadsheet. They compare a one-time GPU price against a monthly API bill and call it a win. But a card pulling 300W around the clock costs real money every month, and community reports from people actually running 24/7 rigs land between $46 and $93. Against a provider charging around $0.14 per million input tokens, power alone can lose. Here is the honest math, including the two things the buy-side arguments usually get right and the rent-side ones ignore: usage changes when you own the hardware, and rental dollars have no resale value.
Trying to work out if a local rig pays for itself?
See our AI training options. We'll run the numbers on your actual usage and set up OpenClaw either way, free.
True cost is sticker minus resale, plus watts. A used 3090 has already taken its depreciation hit; a new 5090 has not. Whichever you buy, the power line is the same shape.
Amazon affiliate links — we earn a small commission at no cost to you.
Bottom Line (July 2026)
- Power is not a rounding error. Owners running 24/7 rigs report electricity bills in the $46 (low) to $93 (high) per month range. That is the line item missing from most “local is free” posts.
- On pure inference cost, the API usually wins. The community line is blunt: power alone is likely to cost more than API access for the same model, because cheap providers price open-weight models near $0.14 per million input tokens.
- Undervolting is the biggest single lever. One report: 320W → 160W at roughly the same usable throughput, via a core-clock cap around 1GHz using LACT.
- True hardware cost is sticker minus resale. Rental dollars are gone; a used 3090 still holds real value years later. That asymmetry is the strongest honest argument for buying.
- Ownership changes usage. A renter disciplined at 8–10 hrs/month became a 60+ hrs/month user after buying. Whether that is a benefit or a rationalization depends on what you do with the hours.
- The consensus advice for most people: max out your Claude and ChatGPT subscriptions first, before spending $20K on a rig.
The Formula Nobody Runs
Almost every “is local cheaper” comparison stops at one-time GPU price versus monthly API bill. That comparison is wrong in both directions: it overstates the hardware cost by ignoring resale, and understates it by ignoring power.
The honest version has three terms:
Monthly cost of local = (hardware price − expected resale) / months you keep it + (avg watts / 1000) × hours per day × 30 × your $/kWh + whatever you still pay for a hosted model anyway
That last term is the one people delete from the spreadsheet and keep paying in real life. Most local-rig owners still hold a Claude or ChatGPT subscription for the hard problems.
Compare that total against your actual current spend, not your imagined spend. Pull the real number off your API dashboard.
Step 1: The Power Line
Electricity is the term you can compute exactly, so start there. All you need is average draw under load, hours per day, and your utility rate.
Worked example with clearly stated assumptions: 300W average draw (a single 24GB card under sustained inference plus the rest of the system), $0.15/kWh.
| Usage pattern | Hours/day | kWh/month | Cost @ $0.15/kWh | Cost @ $0.30/kWh |
|---|---|---|---|---|
| Evening tinkering | 2 | 18 | $2.70 | $5.40 |
| Workday use | 8 | 72 | $10.80 | $21.60 |
| Always-on agent | 24 | 216 | $32.40 | $64.80 |
| Always-on, dual-GPU (600W) | 24 | 432 | $64.80 | $129.60 |
Those computed figures line up with what 24/7 owners actually report: $46 at the low end, $93 at the high end, which is what you get once you add a second card, a hungry CPU, or a region with expensive power. Check your own bill for the $/kWh rather than using ours.
The number that should stop you: at roughly $0.14 per million input tokens from a cheap hosted provider, $46 of electricity buys around 300 million input tokens of API. Almost nobody personally consumes 300M tokens a month. That is the whole reason the community line reads “power alone is likely to cost more than API access for the same model.”
Step 2: Undervolting Cuts the Power Line in Half
If you are keeping the rig anyway, this is the highest-leverage thing on the page.
Local inference is memory-bandwidth bound, not core-clock bound. The card spends most of its time waiting on VRAM, so capping the core clock costs far less throughput than it saves in watts. One community report using LACT on Linux with a core clock capped near 1GHz took a card from about 320W to about 160W with roughly the same usable generation speed.
Halve the watts and you halve the electricity line:
| Setting | Avg draw | kWh/month (24/7) | Cost @ $0.15/kWh |
|---|---|---|---|
| Stock | 320 W | 230 | $34.56 |
| Core clock capped ~1GHz | 160 W | 115 | $17.28 |
That is roughly $200 a year, for a config change, with little practical speed loss. It also runs cooler and quieter, which matters if the machine lives in the room you work in.
Two caveats worth stating plainly. This is one person’s report on their card, not a benchmark we ran, and results vary by GPU generation and by model. Measure your own tokens/sec before and after rather than assuming the trade is free. Prompt processing is more compute-bound than generation, so a heavy prefill workload will feel the clock cap more than a chat workload will.
Step 3: The Rent-vs-Buy Holes
A widely-upvoted breakdown of this argument points at three things the “just rent GPUs, it’s cheaper” math quietly omits.
1. Ownership changes usage. Someone disciplined enough to rent only 8–10 hours a month became a 60+ hour a month user once they owned the card. This is not a footnote, it is the main event. The rental math assumed a usage level that only existed because the meter was running. Whether the extra 50 hours are worth anything depends on what you actually do with them, but for anyone learning, fine-tuning, or building, unmetered time has real value that the rental spreadsheet prices at zero.
2. Rental dollars are gone. Every dollar spent renting has no residual value. Hardware does. A used 3090 bought today still commands a real price years from now, so the honest cost of ownership is sticker minus resale, not sticker. That single correction moves the break-even point substantially in favor of buying.
3. Cards age slowly for this workload. A four-year-old card still runs current models. Local LLM inference cares about VRAM capacity and memory bandwidth, and neither has improved as fast as marketing cycles suggest. The 3090 is the standing proof: still a legitimate recommendation years after release.
Worked Break-Even Example
Assumptions stated up front so you can swap in your own: used RTX 3090 at $800, expected resale $400 after three years, 300W average, 24/7 operation, $0.15/kWh.
| Line item | Math | Per month |
|---|---|---|
| Depreciation | ($800 − $400) / 36 months | $11.11 |
| Electricity, stock | 216 kWh × $0.15 | $32.40 |
| Total, stock | — | $43.51 |
| Total, undervolted (160W) | $11.11 + $17.28 | $28.39 |
| Total, undervolted + 8 hrs/day | $11.11 + $5.76 | $16.87 |
Now compare honestly. Against a $20/month subscription, the always-on stock rig never breaks even on cost. Undervolt it and run it only when you use it, and it lands at roughly the same price as the subscription, with a card you still own at the end.
The interesting rows are not the top ones. They are the last two: the rig only becomes cost-competitive when you stop leaving it running at full clocks out of habit.
And if you scale the sticker price up, the argument collapses fast. A $20K workstation depreciating over three years is $555/month before a single watt. That is why the community advice is what it is: max out your Claude and ChatGPT subscriptions first, then find out whether you actually hit a wall those subscriptions could not solve.
Want the numbers for your exact setup?
The token speed & cost estimator compares a rig against cloud API spend, and the local LLM calculator shows which models your VRAM actually holds — because a card that can't run the model you want has infinite cost per token.
The Honest Split: What Local Actually Wins
Cost per token is the wrong hill for local inference to die on. Here is the split as the community actually reports it.
The API usually wins on:
- Pure inference cost. Especially against cheap providers pricing open-weight models near $0.14/M input.
- Frontier quality. Nothing you run on one consumer card matches the best hosted models on hard problems.
- Zero idle cost. You pay for tokens, not for a machine sitting at idle draw all night.
- No capital risk. No $2,000 bet that a card will still be the right card in two years.
Local wins on:
- Privacy and sovereignty. Your code, your client data, your prompts never leave the building. For regulated work this is not a preference, it is the requirement, and no price comparison applies.
- Unmetered tinkering. The 8-to-60-hour usage jump is the point. No per-token anxiety means you actually experiment.
- Training and fine-tuning. Iterating on a fine-tune against a rented meter is a different psychological experience than iterating on hardware you own.
- Availability. No rate limits, no provider deprecating your model, no outage, no terms-of-service change.
- Batch and background work. An agent that runs all night on your own hardware has a fixed cost. The same agent on an API has an unbounded one.
Notice that none of the local wins are “it’s cheaper.” If cost is your only reason for going local, the math above is telling you something and it is worth listening to.
Common Mistakes
- Leaving out the power bill entirely. It is $30–$90 a month on an always-on rig, every month, forever. That is the single most common omission.
- Comparing sticker price to API cost. Subtract expected resale. Rental and API dollars have no residual value; hardware does.
- Running at stock clocks 24/7. Inference is bandwidth-bound. A core clock cap can roughly halve draw at similar usable throughput. Test it on your card.
- Pretending you will cancel the subscriptions. Most local rig owners keep paying for a hosted model too. Put that line back in the spreadsheet.
- Justifying a $20K build on cost savings. At $555/month of depreciation alone it is not a savings play. Buy it for privacy, capability, or because you want it, but be honest about which.
- Using imagined API spend instead of real API spend. Open your billing dashboard. The number is usually lower than the number in your head.
- Ignoring your actual utility rate. The gap between $0.10/kWh and $0.35/kWh regions changes the entire conclusion.
Related Guides
- Is a $5K local AI rig worth it? — the same question at a bigger sticker price
- RTX 5090 vs 4090 vs used 3090 — which card holds its value
- Mac Studio vs RTX workstation — unified memory draws far fewer watts
- OpenClaw API costs compared — the other side of the ledger
- Zero-dollar OpenClaw with local models — “free” is never quite free
- Hermes agent real token cost — what an agent loop actually consumes
Need OpenClaw fixed live?
Remote rescue sessions for gateway, auth, tunnel, VPS, and model access problems.
See Rescue Session