Includes $31.03 electricity at 170 W
Tools / Cost of local inference
The full cost of useful work on your own AI workstation: machine, setup, software, upkeep and electricity over your horizon, beside your API bill. Every input is editable; nothing is hidden in a multiplier.
Planning scenario · 36-month comparison · editable assumptions
Estimated cloud advantage · 99% below the local total
WHERE LOCAL STARTS TO PAY
That is about 2,373 tokens per second, input and output combined, around the clock, at gpt-oss-120b, hosted prices and your input/output mix. You entered 60 million, 1% of that threshold.
Includes $31.03 electricity at 170 W
Cloud monthly cost does not exceed local monthly cost
AT YOUR ENTERED DEMAND
Local cost is the total over the horizon spread across the months and your monthly demand, so unused capacity is charged to the work actually delivered. Planning ratios, not measured hardware performance, and no claim of equal model quality.
Local = hardware + setup + months × (software and support + administration + electricity) − resale value at the end. Electricity = watts ÷ 1,000 × powered hours × tariff. Cloud = input millions × input price + output millions × output price (default: dated hosted prices for the same open model), or a flat monthly bill. Break-even volume = local monthly cost (upfront minus resale, spread over the horizon, plus running costs) ÷ the blended cloud price per million tokens. Payback = upfront ÷ (cloud monthly − local monthly) when positive. The chart subtracts resale value only at its final point.
Prefilled hardware prices are dated regional offers on the seller’s tax basis; prefilled power is 55% of the highest rating in the machine record, a placeholder for a wall-meter reading. No financing, tax or cooling multiplier is applied. Method.