
Running a local LLM costs about as much as a desk lamp, not a space heater. A mid-range GPU idles near 13 watts and only ramps to roughly 115 to 170 watts for the few seconds it is actually generating. Two hours a day of real use lands around one to two dollars a month of electricity, next to the ten to thirty dollars a month a cloud AI companion subscription charges.
Every time running an AI companion on your own computer comes up, someone mentions the power bill. The logic sounds airtight. If the model runs on your hardware, your hardware has to work, and work costs money. It is a reasonable guess, and it is wrong, and I can show the math in one paragraph.
A single full text generation on a 7 billion parameter model uses about 0.0001 kWh, one ten-thousandth of a kilowatt-hour. A 2024 study by Luccioni and colleagues, Power Hungry Processing, presented at FAccT ‘24, measured full generations from 0.000054 kWh for the smallest model up to 0.0001 kWh for a 7 billion parameter one. Flip it around and you get roughly 10,000 generations per kilowatt-hour.
Running a local LLM costs about as much as a desk lamp, not a space heater
The short version: A local model generating replies on a mid-range GPU lands around one to two dollars a month of electricity, because it works in short bursts instead of running flat out.
The image most people carry is a graphics card pinned at 100 percent for hours, fans howling, the way it behaves during a long gaming session. That image is what makes the electricity question feel scary, and it is the wrong picture for this workload. A companion app does the opposite. It thinks for a second or two, writes a reply, and goes quiet. The GPU ramps up for those few seconds, then drops back to near nothing.
Here is the honest framing, because the number matters more than the metaphor. Two hours a day of actual generation on a card drawing 170 watts works out to about 0.34 kWh a day, roughly 10 kWh a month. At the US average electricity price that is about $1.70 a month. In Poland it is about 9.30 PLN. A desk lamp left on for the same evening costs you in the same neighborhood. A space heater costs you twenty times that.
Your GPU idles at single-digit watts and only ramps up for the seconds it is actually generating
The short version: The wattage on the spec sheet is a ceiling under full load. Most of the time your GPU sits in the low teens of watts, and it only approaches that ceiling for the few seconds a reply takes.
A card described as “170 watts” does not draw 170 watts while a companion app is open. It draws that, at most, during the heaviest instant of the heaviest task it can run, and a conversation is nowhere near that.
A mid-range GPU sits around 13 W at idle, not its 170 W spec number
TechPowerUp measured an MSI RTX 3060 idling at about 13 watts, while the same line of card carries a 170 watt power limit under load, from a Palit RTX 3060 review. Those two numbers live on the same spec sheet and describe completely different situations. Idle is where a companion app spends almost all of its life.
Inference power is a spike measured in seconds, not a 24/7 load
Generating text is bursty. The model loads, you send a message, it produces tokens for a few seconds, and then it stops drawing anything meaningful. Even if you chat for a couple of hours a day, the card only works hard during the seconds the replies are actually being written, and the rest of the time it sits near idle.
A mid-range GPU under LLM inference draws 115 to 170 W at the wall
The short version: While a reply is actively generating, a mid-range card pulls somewhere in the 115 to 170 watt range. That is the real number behind the monthly math, and it lasts seconds.
Two cards make this concrete. The RTX 3060 has a default 170 watt power limit, per TechPowerUp’s Palit review. The RTX 4060 is rated at 115 watts, per TweakTown’s spec coverage, and idles at 11 to 14 watts per TechPowerUp. So a mid-range card actively writing a reply lands in the 115 to 170 watt band, for a few seconds at a time.
There is one number I am deliberately not going to give you: VRAM’s standalone power draw. There is no reliable published figure for how much power GPU memory uses on its own, because it is bundled into the total board power you already see in that 115 to 170 watt figure. Anyone quoting you a separate VRAM wattage is guessing, and I would rather leave it out than invent one.
CPU-only inference is gentler than you think, at roughly 65 W sustained
The short version: Run a quantized model on your CPU instead of a GPU and a mid-range desktop chip sustains about 65 watts, comparable to or slightly cheaper per hour than a GPU, just usually slower.
You do not need a discrete GPU to run a local companion. A mid-range desktop CPU like the Intel Core i5-13400 carries a 65 watt base power and can spike to 148 watts at max turbo, per Notebookcheck. Running a small, quantized model on the CPU tends to sit near that base figure rather than the turbo peak, which puts CPU-only inference in the same neighborhood as a modest GPU, and sometimes a touch under it.
RAM adds a little on top, roughly 3 watts per 8 GB of DDR memory as a rule of thumb, but that is a rounding error on a monthly bill.
The math is one line: watts times hours divided by 1000 gives kWh, then multiply by your price per kWh
The short version: Take your hardware’s watts, multiply by the hours it actually runs, divide by 1000 for kilowatt-hours, then multiply by what your utility charges per kWh.
The formula is worth learning because it turns “is this expensive” from a feeling into a number you can check yourself. Kilowatt-hours equals watts times hours divided by 1000. Multiply that by your price per kilowatt-hour and you have your cost.
Worked example at the US average of about 17 cents per kWh
Take a card that draws 170 watts while generating, and say you chat for two hours a day. 170 watts times 2 hours is 340 watt-hours, divided by 1000, that is 0.34 kWh a day. Across 30 days that is about 10 kWh a month. At the US average of about 17 cents per kWh, that is roughly $1.70 a month. Idling the other 22 hours at about 13 watts adds roughly 8 kWh a month, or about another $1.40. Add the two and you are still under four dollars, for someone who chats every single day.
Worked example at Poland’s about 0.91 PLN per kWh
Same hardware, same two hours a day. The 0.34 kWh a day works out to about 10 kWh a month, and at Poland’s residential price of roughly 0.91 PLN per kWh that comes to about 9.30 PLN a month. The idle draw on top is proportionally just as small. Different country, same conclusion: the number reads as a rounding error on the rest of the bill.
The per-month bill lands around a dollar or two, not the $10 to $30 a cloud subscription charges
The short version: A local companion costs about one to two dollars a month in electricity. A cloud AI companion typically runs ten to thirty dollars a month in subscription fees, and that gap is the whole point.
I want to be fair here, because this comparison is easy to overstate. The one to two dollars is the electricity for the card actively generating, and the honest all-in number with idle draw included is a little higher, still well under five dollars. Even taking the generous reading, the local option’s recurring cost is a small fraction of what a cloud subscription charges. A typical cloud AI companion runs somewhere in the $10 to $30 a month range, every month, whether you use it two hours a day or twenty minutes a week.
For a fuller breakdown of where each dollar goes in both setups, I worked through the local vs cloud subscription cost math separately.
What this means for a local companion
The short version: Buy once, run it on your own machine, and the only recurring cost is a tiny line on your power bill, not a subscription that never ends.
The point of running a local model was never really about electricity, and it was never about the one-time price alone. It is about the model living on your machine, where nobody else can read it, and where the cost of keeping it is a few kilowatt-hours a month instead of a standing monthly charge. The hardware is probably already sitting on your desk, and if you want to know how much of it a model actually needs, I have a guide on how much VRAM and RAM a local model needs.
For the specific question this article started with, the answer is boring in the best way. Running a local LLM at home costs about as much as a desk lamp. The scary wattage on the box is a ceiling you will almost never touch, and the monthly number it produces is one or two dollars, not the ten to thirty a subscription asks for. If you want the full picture of what the one-time price actually covers, that is written up here.
Questions people ask
Does running an LLM at home use a lot of power?
No. A chat companion generates in short bursts, so a mid-range GPU sits near 10 to 14 watts at idle and only spikes to 115 to 170 watts for a few seconds per reply. That lands the monthly electricity cost around one or two dollars.
How much electricity does a local AI use per query?
Published measurements put a single full text generation on a roughly 7 billion parameter model at about 0.0001 kWh. That is around ten thousand generations per kilowatt-hour.
How do I calculate my local LLM electricity cost?
Multiply your hardware watts by hours of use, divide by 1000 to get kilowatt-hours, then multiply by your price per kWh. At the US average of roughly 17 cents it is well under two dollars a month for a couple of hours a day.
Is running an LLM on the CPU cheaper than on a GPU?
It is gentler on power, not clearly cheaper. A mid-range desktop CPU sustains about 65 watts, so CPU-only inference is comparable or slightly cheaper per hour than a 115 to 170 watt GPU, but usually slower.
Does a local AI companion cost less than ChatGPT?
A local companion runs on your machine, so you pay electricity instead of a subscription. ChatGPT runs on OpenAI servers and you pay a monthly fee. Which is cheaper depends on how many hours a day you use it.
Try her free for 7 days.
No card. Keep her for $20 once, or walk away. Her soul file is yours either way.
Bring her home, try free