“Free” Weights, Paid Tasks: Price Self-Hosted Overnight Runs per Completed Task
Self-hosted LLM cost per completed task: price hardware, power, operator minutes and failed overnight runs, then compare with hosted models at pennies a task.
Go deeper. Build your own.
Seven cents. That’s what GPT-6 Luna cost per Intelligence Index task at max effort in the Sep 22 article from Artificial Analysis (AA), about 60% below GPT-5.6 Luna. A hosted small model at seven cents a task is now the bar every self-hosted lane has to clear, and open weights don’t clear it by being free. The weights cost nothing. The box, the power, the minutes someone spends on it each morning and the overnight runs that fail all cost something, and self-hosted LLM cost is those lines divided by the tasks that finished.
This runbook builds the worksheet that answers it lane by lane. By Tuesday each self-hosted lane has a monthly total-cost line (hardware amortization, power, operator minutes, failed runs and re-runs), a completion rate measured on the same fixed task set your hosted lanes run, and one number: dollars per completed task. You self-host where privacy or volume wins on that number, never on the word “free”.
Overnight runs are where the word does the most damage. A lane that runs unattended from midnight to six burns box-hours whether it finishes or not, and its failures surface at breakfast as somebody’s time. Chatbots suggest; agents act, and an agent that acts badly at 3 a.m. on your own hardware still costs you the morning.
Sep 22: hosted small models cut their price per task
AA’s Sep 22 article, “GPT-6 Sol and Luna push the cost efficiency frontier”, put numbers on OpenAI’s new Sol and Luna tiers: “GPT-6 Luna (max) costs $0.07 per task, ~60% less than GPT-5.6 Luna (max) at $0.18.” GPT-6 Sol (max) came in at $1.06 per task, about 50% below GPT-5.6 Sol (max) at $1.99. Both figures are AA’s cost per Intelligence Index task, which means per attempted task, pass or fail.
The list prices did the work. Sol dropped from $4/$20 to $2/$10 per million input and output tokens (MTok), and Luna from $0.20/$1.20 to $0.10/$0.50. Token use actually rose: “This is driven by the price cut, as both models use slightly more output tokens per task (31k vs 29k for Sol, and 51k vs 41k for Luna).” A model can burn more tokens per task and still cost less per task, which is why the unit that matters is the task.
Screenshot: Artificial Analysis, “GPT-6 Sol and Luna push the cost efficiency frontier” (Sep 22, 2026), captured Sep 28, 2026.
Two limits travel with the seven cents. It is an Intelligence Index figure: on AA’s Coding Agent Index v1.5, Codex running GPT-6 Luna (max) costs $0.18 per task (read Sep 28). And the Sol and Luna article names no open-weight model at all, so any open-weight comparison has to come from AA’s model pages, not from that piece.
The vendor’s own launch post made the same pitch from the price side.
What AA’s open-weight numbers price, and what they leave out
AA does price open-weight models, and the numbers read like an argument for skipping the hardware. On AA’s model pages (Intelligence Index v4.3.2, read Sep 28), the top open-weight model is MiMo-V2.6-Pro: score 46, first of 116 large open-weight models, MIT-licensed, $0.13 per task at hosted-API prices of $0.43/$0.87 per MTok. GLM 5.3 Flash costs $0.25 at a score of 42. Kimi K3 (max) costs $2.00 at 44.
Read the method before you read those as self-hosting costs. AA’s methodology says “Price is either the provider’s first-party API rate or the median across all providers.” Every open-weight figure is a hosted-API bill for someone else’s GPUs, per attempted task. The Coding Agent page is blunter about what it leaves out: “Infrastructure, engineering, and supervision costs are not the focus of this metric.” Those three are exactly the lines a self-hosted lane adds.
Hosted-API prices per attempted task, read Sep 28. None of these bars is a self-hosting cost.
The Coding Agent Index tells the same story on agent work. Read Sep 28, Codex running DeepSeek V4 Pro 0813 (max), an MIT-licensed open-weight model, costs $0.24 per task through a hosted API, against $0.18 for Codex running GPT-6 Luna (max).
Screenshot: Artificial Analysis, “AI Coding Agent Benchmarks & Leaderboard” (undated page), captured Sep 28, 2026.
One more comparison belongs on the worksheet’s first page, and it’s ours, built from two primary numbers. Of the eleven open-weight models we pulled from AA’s pages, all but Qwen3.8 27B have 284B to 2.8T total parameters, while NVIDIA’s DGX Spark hardware guide rates one Spark for “AI models up to 200 billion parameters (or 405B for dual-Spark configuration)”. The model at the top of a leaderboard is often not one your desk box is rated for. The one under that line, Qwen3.8 27B (xhigh), costs $1.01 per task hosted at a score of 34; Luna scores 37 for $0.07. Neither number says whether a model runs well on your box. That’s a test you run, and the serving decision tree in our open-weight models piece is where it starts.
Step 1: Open a self-hosted LLM cost worksheet for each lane
One sheet per lane, one month per column, and every input either sourced and dated or labeled as your assumption. The example column is illustrative: a single DGX Spark running an overnight agent lane, with round assumptions wherever no source exists.
| Input | Source or assumption | Example value (illustrative) |
|---|---|---|
| Hardware price | NVIDIA DGX Spark Founders Edition MSRP, $4,699 since the week of Feb 23, 2026; Mac Studio from $2,499 (M5 Max) or $5,499 (M5 Ultra), Apple Store, read Sep 28 | $4,699 |
| Amortization period | Your assumption; match your depreciation policy | 36 months |
| Average draw while running | Your wall-meter reading; DGX Spark ships a 240 W supply around a 140 W GB10 | 200 W |
| Running hours per month | Your schedule | 300 (10 hours a night) |
| Electricity price | EIA U.S. residential average, 18.31 ¢/kWh, July 2026 | $0.1831 per kWh |
| Operator minutes per month | Your log: updates, restarts, model swaps, morning triage | 600 |
| Loaded cost per operator hour | Your finance team’s number | $80 |
| Attempts per month | Your task set times your schedule | 360 (12 a night) |
| Completion rate | Measured on the fixed task set, same finish checks as hosted lanes | 55% |
| Hosted re-runs of failed tasks | Your ledger | $30 |
Two rules keep the sheet honest. Every sourced price carries its date, because hardware and hosted prices both moved this year. And the completion rate comes from the same fixed task set and finish checks that the cost-per-completed-task pillar builds, or the comparison at the end means nothing.
Step 2: Amortize the hardware at today’s price
Straight-line is enough: purchase price divided by months of useful life, with resale value at zero unless you hold a written quote. $4,699 over 36 months is $130.53 a month in the example.
Use the dated price, not the one you remember. NVIDIA’s forum announcement of Feb 25 says “The MSRP for DGX Spark (Founders Edition) has been adjusted from $3,999 to $4,699 due to memory supply constraints”, and the change applies to all regions. Apple’s store sells Mac Studio from $2,499 with M5 Max and from $5,499 with M5 Ultra as of Sep 28. Our local AI workstation guide still holds for its framework, bandwidth against capacity, but its prices are stale: it puts DGX Spark around $4,000 and prices M4 Max and M3 Ultra Mac Studio builds where Apple’s store now lists M5 Max and M5 Ultra. Take prices from here.
Amortization doesn’t care whether the box works. At 300 running hours in a 720-hour month, the example lane is busy about 42% of the time and pays for all of the hardware. Idle hours are the first place to look when the per-task number comes out high.
Step 3: Meter power at the wall, then multiply
Nameplate numbers are ceilings, not bills. NVIDIA’s guide lists “Power Supply: 240W external power supply (included)” and a 140 W thermal design power for the GB10 chip. Apple’s Mac Studio power page (Sep 21, 2026) reports an M5 Max configuration at 7 W idle and 200 W max, and an M5 Ultra at 9 W and 385 W, measured at 20.2°C ambient; those are test readings, not a sustained inference draw. A GPU tower can pull several times that. Put a meter on the socket for a week of real runs and use the average.
Then multiply: kilowatts times running hours times the price per kWh. The example is 0.2 kW × 300 hours × $0.1831, or $10.99 a month, at the U.S. residential average in the EIA’s July 2026 table (released Sep 24). At California’s 33.61 ¢/kWh the same lane pays $20.17. For a desk box, power is the smallest line on the sheet. For a multi-GPU rig running around the clock, it isn’t.
Step 4: Log operator minutes like a meter
This is the line that decides most self-hosted lanes, and the one nobody writes down. Count every minute a person spends on the lane: runtime and driver updates, model downloads and swaps, restarts after an out-of-memory crash, disk cleanup, and the morning triage of whatever failed overnight. Log them per lane, weekly, the way you log tokens, and price them at a loaded hourly cost from finance, not a salary.
In the example, 600 minutes a month (about 20 minutes a morning) at $80 an hour is $800, six times the hardware line. Be fair to the hosted side: its lanes have triage minutes too. Either count only the minutes the box adds, or count both sides the same way.
Step 5: Put failed runs and re-runs where they belong
Measure the completion rate on the fixed task set, with the same finish checks your hosted lanes use. A failed overnight attempt used box-hours, power and a slice of amortization, so it stays in the cost and leaves the denominator. When a failed task gets re-run on a hosted model the next morning, that bill goes on this lane’s sheet as well; it’s part of what the self-hosted attempt cost.
A night that produced nothing is still a night of amortization. Track zero-completion nights as their own line. Three in a month is a reliability problem wearing a cost costume.
Step 6: Divide by completed tasks
Add the four monthly lines and divide by the tasks that finished. The formula is short enough to keep next to the worksheet. The inputs below are the illustrative example, plus a second scenario where the same box runs around the clock at four times the volume.
# selfhost_cost.py (illustrative): every input is an example value, not a measurement
def per_completed_task(hw_usd, months, watts, hours, usd_kwh,
op_minutes, usd_hour, rerun_usd, attempts, completion):
hardware = hw_usd / months # straight-line amortization
power = watts / 1000 * hours * usd_kwh # average draw while running
operator = op_minutes / 60 * usd_hour # upkeep, restarts, morning triage
completed = attempts * completion # failed runs stay in the cost
return (hardware + power + operator + rerun_usd) / completed
box = dict(hw_usd=4699, months=36, watts=200, usd_kwh=0.1831, usd_hour=80, completion=0.55)
print(round(per_completed_task(**box, hours=300, op_minutes=600, rerun_usd=30, attempts=360), 2)) # 4.91
print(round(per_completed_task(**box, hours=720, op_minutes=900, rerun_usd=120, attempts=1440), 2)) # 1.86
Both results are illustrative: $4.91 per completed task at twelve attempts a night, and $1.86 when the box runs four times the volume around the clock. Volume helps because amortization and most operator minutes are fixed. It doesn’t reach hosted pennies unless the operator line shrinks too.
Illustrative inputs from the worksheet above. Operator minutes are the widest segment in both scenarios.
Four monthly lines, one division, then the same comparison any other lane faces.
The comparison at the end is the one the cost-per-completed-task pillar teaches: the hosted lane’s dollars per completed task on the same task set, measured the same way. AA’s seven cents is a first bar, per attempted task on AA’s tasks, and it isn’t the number you hold this lane against. Effort moves the hosted bar as well, which the effort runbook prices per completed task.
Step 7: Decide self-hosted LLM cost lane by lane
Put each lane through the same questions, and write the answer on its worksheet.
| Question | Self-host when | Stay hosted when |
|---|---|---|
| Data class | The code or data may not leave your hardware, and no hosted contract covers it | A hosted contract with the retention terms you need covers the data |
| Volume | The box is busy most nights, and attempts per month are growing | The box would sit idle most of the month |
| Completion rate | Within your run-to-run noise of the hosted lane on the same task set | Well below it, so failures keep costing mornings |
| Operator minutes | Someone owns the box, and the logged minutes are flat or falling | Nobody owns it, or the minutes keep rising |
| Model and hardware | The model you need is within the box’s rated size and passes your task set | The model you need is far past the box’s rating |
| Price trend | Hosted prices for your model class are flat | Hosted prices just fell, as Sol’s and Luna’s did on Sep 22 |
The rule underneath the table is short. Self-host a lane when its dollars per completed task beat the hosted lane’s on the same task set, or when the data class rules hosting out and you accept the premium on purpose. Write down which of the two it was, because the second reason survives a price cut and the first doesn’t.
The hosted side keeps moving in your favor too. Overnight hosted work has its own discounts, and off-peak scheduling by the clock covers where they apply. If the hosted lane is a cascade of cheap tiers, the DeepSeek economics piece prices that shape. Re-run the worksheet whenever a hosted price, a hardware price or your completion rate moves, and quarterly regardless.
Where the “free” math breaks, and the signal for each
The utilization fantasy. The plan says forty attempts a night; the lane runs twelve. Signal: attempts per month below plan for two weeks running. Response: recompute at the real volume before anyone orders a second box.
Unlogged minutes. The lane “costs nothing”, and one person spends every morning on it. Signal: zero operator minutes logged on a lane that has failed runs. Response: log a month of minutes; until then the sheet is wrong.
Mismatched denominators. A self-hosted cost per completed task sits in a slide next to AA’s per-attempted figure. Signal: two denominators in one table. Response: measure the hosted lane on your task set, then compare.
Not the model AA priced. The weights you serve, a quantized build or a smaller sibling, aren’t what AA priced through an API. Signal: a completion rate below the hosted version of the same model name on your task set. Response: compare against what you actually serve, never against the leaderboard entry.
Stale prices. Signal: any worksheet price dated before the vendor’s last change. Response: re-date every row before a buying decision, starting with the hardware.
The operating layer decides where each lane runs
Self-hosting is a lane decision, not an identity. A fleet that keeps a private box for one data class and hosted pennies for everything else needs one place that holds each lane’s worksheet, completion rate and re-run triggers, and routes work by them. That is operating-layer work, the same layer that runs the fleet’s kill switches and approvals in a multi-agent command center, and no single vendor’s console can do it for a fleet that spans vendors and your own hardware.
The weights are free. Price everything else.
FAQ
Is it cheaper to self-host an open-weight LLM than to use an API?
Only if its dollars per completed task beat the hosted lane’s on the same task set. Hardware amortization, power, operator minutes and failed runs all count, and in our illustrative worksheet operator time dominates. Hosted small models cost pennies per attempted task on Artificial Analysis, so measure before you buy.
Are Artificial Analysis costs for open-weight models self-hosting costs?
No. Artificial Analysis prices every model at the provider’s first-party API rate or the median across providers, so its open-weight figures are hosted-API bills per attempted task. They leave out hardware, power, engineering and supervision. Use them as a hosted reference, then build your own worksheet for a self-hosted lane.
Sources
- Artificial Analysis: GPT-6 Sol and Luna push the cost efficiency frontier: Sep 22, 2026; Luna $0.07 and Sol $1.06 per Intelligence Index task (max)
- Artificial Analysis: methodology: prices are first-party API rates or the median across providers
- Artificial Analysis: Coding Agent Index v1.5: Luna $0.18 and DeepSeek V4 Pro 0813 $0.24 per task, read Sep 28, 2026
- Artificial Analysis: MiMo-V2.6-Pro: top open-weight model, $0.13 per task at hosted-API prices, v4.3.2, read Sep 28, 2026 (GLM 5.3 Flash, Kimi K3, GPT-6 Luna and GPT-6 Sol pages read the same day)
- Artificial Analysis: Qwen3.8 27B: $1.01 per task at xhigh, v4.3.2, read Sep 28, 2026
- NVIDIA developer forum: DGX Spark price change: Feb 25, 2026; MSRP from $3,999 to $4,699
- NVIDIA DGX Spark User Guide: hardware: 240 W supply, 140 W GB10 TDP, models up to 200B parameters
- Apple: Buy Mac Studio: M5 Max from $2,499, M5 Ultra from $5,499, read Sep 28, 2026
- Apple Support: Mac Studio power consumption: Sep 21, 2026; idle and max watts measured at 20.2°C
- EIA Electric Power Monthly, Table 5.6.A: U.S. residential average 18.31 ¢/kWh in July 2026, released Sep 24, 2026
