← All guides

Local LLM Electricity Cost vs API (July 2026): The Break-Even Math Including Power and Depreciation

Nobody puts the power bill in the spreadsheet. They compare a one-time GPU price against a monthly API bill and call it a win. But a card pulling 300W around the clock costs real money every month, and community reports from people actually running 24/7 rigs land between $46 and $93. Against a provider charging around $0.14 per million input tokens, power alone can lose. Here is the honest math, including the two things the buy-side arguments usually get right and the rent-side ones ignore: usage changes when you own the hardware, and rental dollars have no resale value.

Trying to work out if a local rig pays for itself?

See our AI training options. We'll run the numbers on your actual usage and set up OpenClaw either way, free.

🎮 CARDS, RANKED BY COST TO OWN

True cost is sticker minus resale, plus watts. A used 3090 has already taken its depreciation hit; a new 5090 has not. Whichever you buy, the power line is the same shape.

Amazon affiliate links — we earn a small commission at no cost to you.

Bottom Line (July 2026)

  • Power is not a rounding error. Owners running 24/7 rigs report electricity bills in the $46 (low) to $93 (high) per month range. That is the line item missing from most “local is free” posts.
  • On pure inference cost, the API usually wins. The community line is blunt: power alone is likely to cost more than API access for the same model, because cheap providers price open-weight models near $0.14 per million input tokens.
  • Undervolting is the biggest single lever. One report: 320W → 160W at roughly the same usable throughput, via a core-clock cap around 1GHz using LACT.
  • True hardware cost is sticker minus resale. Rental dollars are gone; a used 3090 still holds real value years later. That asymmetry is the strongest honest argument for buying.
  • Ownership changes usage. A renter disciplined at 8–10 hrs/month became a 60+ hrs/month user after buying. Whether that is a benefit or a rationalization depends on what you do with the hours.
  • The consensus advice for most people: max out your Claude and ChatGPT subscriptions first, before spending $20K on a rig.

The Formula Nobody Runs

Almost every “is local cheaper” comparison stops at one-time GPU price versus monthly API bill. That comparison is wrong in both directions: it overstates the hardware cost by ignoring resale, and understates it by ignoring power.

The honest version has three terms:

Monthly cost of local
  = (hardware price − expected resale) / months you keep it
  + (avg watts / 1000) × hours per day × 30 × your $/kWh
  + whatever you still pay for a hosted model anyway

That last term is the one people delete from the spreadsheet and keep paying in real life. Most local-rig owners still hold a Claude or ChatGPT subscription for the hard problems.

Compare that total against your actual current spend, not your imagined spend. Pull the real number off your API dashboard.

Step 1: The Power Line

Electricity is the term you can compute exactly, so start there. All you need is average draw under load, hours per day, and your utility rate.

Worked example with clearly stated assumptions: 300W average draw (a single 24GB card under sustained inference plus the rest of the system), $0.15/kWh.

Usage patternHours/daykWh/monthCost @ $0.15/kWhCost @ $0.30/kWh
Evening tinkering218$2.70$5.40
Workday use872$10.80$21.60
Always-on agent24216$32.40$64.80
Always-on, dual-GPU (600W)24432$64.80$129.60

Those computed figures line up with what 24/7 owners actually report: $46 at the low end, $93 at the high end, which is what you get once you add a second card, a hungry CPU, or a region with expensive power. Check your own bill for the $/kWh rather than using ours.

The number that should stop you: at roughly $0.14 per million input tokens from a cheap hosted provider, $46 of electricity buys around 300 million input tokens of API. Almost nobody personally consumes 300M tokens a month. That is the whole reason the community line reads “power alone is likely to cost more than API access for the same model.”

Step 2: Undervolting Cuts the Power Line in Half

If you are keeping the rig anyway, this is the highest-leverage thing on the page.

Local inference is memory-bandwidth bound, not core-clock bound. The card spends most of its time waiting on VRAM, so capping the core clock costs far less throughput than it saves in watts. One community report using LACT on Linux with a core clock capped near 1GHz took a card from about 320W to about 160W with roughly the same usable generation speed.

Halve the watts and you halve the electricity line:

SettingAvg drawkWh/month (24/7)Cost @ $0.15/kWh
Stock320 W230$34.56
Core clock capped ~1GHz160 W115$17.28

That is roughly $200 a year, for a config change, with little practical speed loss. It also runs cooler and quieter, which matters if the machine lives in the room you work in.

Two caveats worth stating plainly. This is one person’s report on their card, not a benchmark we ran, and results vary by GPU generation and by model. Measure your own tokens/sec before and after rather than assuming the trade is free. Prompt processing is more compute-bound than generation, so a heavy prefill workload will feel the clock cap more than a chat workload will.

Step 3: The Rent-vs-Buy Holes

A widely-upvoted breakdown of this argument points at three things the “just rent GPUs, it’s cheaper” math quietly omits.

1. Ownership changes usage. Someone disciplined enough to rent only 8–10 hours a month became a 60+ hour a month user once they owned the card. This is not a footnote, it is the main event. The rental math assumed a usage level that only existed because the meter was running. Whether the extra 50 hours are worth anything depends on what you actually do with them, but for anyone learning, fine-tuning, or building, unmetered time has real value that the rental spreadsheet prices at zero.

2. Rental dollars are gone. Every dollar spent renting has no residual value. Hardware does. A used 3090 bought today still commands a real price years from now, so the honest cost of ownership is sticker minus resale, not sticker. That single correction moves the break-even point substantially in favor of buying.

3. Cards age slowly for this workload. A four-year-old card still runs current models. Local LLM inference cares about VRAM capacity and memory bandwidth, and neither has improved as fast as marketing cycles suggest. The 3090 is the standing proof: still a legitimate recommendation years after release.

Worked Break-Even Example

Assumptions stated up front so you can swap in your own: used RTX 3090 at $800, expected resale $400 after three years, 300W average, 24/7 operation, $0.15/kWh.

Line itemMathPer month
Depreciation($800 − $400) / 36 months$11.11
Electricity, stock216 kWh × $0.15$32.40
Total, stock$43.51
Total, undervolted (160W)$11.11 + $17.28$28.39
Total, undervolted + 8 hrs/day$11.11 + $5.76$16.87

Now compare honestly. Against a $20/month subscription, the always-on stock rig never breaks even on cost. Undervolt it and run it only when you use it, and it lands at roughly the same price as the subscription, with a card you still own at the end.

The interesting rows are not the top ones. They are the last two: the rig only becomes cost-competitive when you stop leaving it running at full clocks out of habit.

And if you scale the sticker price up, the argument collapses fast. A $20K workstation depreciating over three years is $555/month before a single watt. That is why the community advice is what it is: max out your Claude and ChatGPT subscriptions first, then find out whether you actually hit a wall those subscriptions could not solve.

Want the numbers for your exact setup?

The token speed & cost estimator compares a rig against cloud API spend, and the local LLM calculator shows which models your VRAM actually holds — because a card that can't run the model you want has infinite cost per token.

The Honest Split: What Local Actually Wins

Cost per token is the wrong hill for local inference to die on. Here is the split as the community actually reports it.

The API usually wins on:

  • Pure inference cost. Especially against cheap providers pricing open-weight models near $0.14/M input.
  • Frontier quality. Nothing you run on one consumer card matches the best hosted models on hard problems.
  • Zero idle cost. You pay for tokens, not for a machine sitting at idle draw all night.
  • No capital risk. No $2,000 bet that a card will still be the right card in two years.

Local wins on:

  • Privacy and sovereignty. Your code, your client data, your prompts never leave the building. For regulated work this is not a preference, it is the requirement, and no price comparison applies.
  • Unmetered tinkering. The 8-to-60-hour usage jump is the point. No per-token anxiety means you actually experiment.
  • Training and fine-tuning. Iterating on a fine-tune against a rented meter is a different psychological experience than iterating on hardware you own.
  • Availability. No rate limits, no provider deprecating your model, no outage, no terms-of-service change.
  • Batch and background work. An agent that runs all night on your own hardware has a fixed cost. The same agent on an API has an unbounded one.

Notice that none of the local wins are “it’s cheaper.” If cost is your only reason for going local, the math above is telling you something and it is worth listening to.

Common Mistakes

  1. Leaving out the power bill entirely. It is $30–$90 a month on an always-on rig, every month, forever. That is the single most common omission.
  2. Comparing sticker price to API cost. Subtract expected resale. Rental and API dollars have no residual value; hardware does.
  3. Running at stock clocks 24/7. Inference is bandwidth-bound. A core clock cap can roughly halve draw at similar usable throughput. Test it on your card.
  4. Pretending you will cancel the subscriptions. Most local rig owners keep paying for a hosted model too. Put that line back in the spreadsheet.
  5. Justifying a $20K build on cost savings. At $555/month of depreciation alone it is not a savings play. Buy it for privacy, capability, or because you want it, but be honest about which.
  6. Using imagined API spend instead of real API spend. Open your billing dashboard. The number is usually lower than the number in your head.
  7. Ignoring your actual utility rate. The gap between $0.10/kWh and $0.35/kWh regions changes the entire conclusion.

Need OpenClaw fixed live?

Remote rescue sessions for gateway, auth, tunnel, VPS, and model access problems.

See Rescue Session

Read next

What Hermes Agent Actually Costs: The Token Bill Nobody Shows You (July 2026)
Tutorials quote the $8-10/mo VPS and stop. Community wire captures show a 40-token 'hi' becoming a 20,538-token request. Here is where the tokens go and the settings people used to cut $15-30/mo down to $2-5.
The $4K Rig That Saves $1K a Week, and the 8-GPU Owner Who Still Uses Claude
I made a video on whether a $5K local AI rig is worth it, and the whole thing comes down to two people. One spent about $4,000 and says he saves a thousand dollars a week with it. The other runs eight graphics cards and still reaches for Claude. Here is the gist, and the framewor
Why OpenClaw Uses 9,600 Tokens for a Simple Question (And How to Fix It)
OpenClaw sends 8,000+ system tokens with every request. Learn where 9,600 tokens go, why costs snowball, and 5 fixes to cut usage fast.
How Much Does OpenClaw Cost Per Month? Full Pricing Breakdown
OpenClaw is free. Your real cost is the API bill: $0-15/month personal, $50-200 business. Full breakdown with cost-cutting tips.