Cost per Completed Task Beats $/MTok: Price the Finished Job, Not the Token
AI agent cost per task beats list price. Count attempts, retries, escalations, cache and tool meters, divide by tasks that finished, and compare models on that.
Go deeper. Build your own.
Read Artificial Analysis’s Coding Agent Index on Sep 28 and one row breaks the way most teams shop for models. Of the three agents at the top, Claude Code running Claude Opus 5.5 has the cheapest token, at $4 in and $20 out per million tokens (MTok), and the most expensive task: $13.0 per attempted task at max effort, against $7.47 for Codex running GPT-6 Astra, whose tokens list at $10 and $50. AI agent cost per task, not price per token, is the number your invoice follows.
That comparison is at max effort, and it counts failed attempts, so it isn’t your number either. This runbook gets you yours. By Tuesday you have a fixed, versioned task set with a finish check on every task; a ledger that counts every attempt, retry, escalation, cache write, cache read and tool or sandbox meter; the effort level each run was actually served at; and one figure per model: dollars per completed task, reported as dollars per merged PR and per finished overnight job. Models get compared on that figure and nothing else.
Chatbots suggest; agents act, and acting is a loop. Every turn resends the conversation and every red test buys another turn. A price sheet can’t see the loop. A ledger can.
Sep 7 to Sep 22: per-task prices went public, and they move
On Sep 7, Artificial Analysis (AA) shipped Intelligence Index v4.3 with a line built for this argument: “GPT-6 Astra (max) and Claude Fable 5.1 (max with fallback) both score 53, but their average cost per Intelligence Index task is $3.26 and $7.63 respectively - 57% lower for Astra.” Both models list at $10/$50 per MTok. Same token price, same score, a 57% gap (max vs max, v4.3, Sep 7).
Three caveats come with that sentence, and all three belong in your method.
It is one effort setting, on one date. Two weeks later the pair stopped being the top of the table: on Sep 22 AA put Claude Opus 5.5 at 58 at max, and at high effort Opus 5.5 scores 54 for $1.82 per Intelligence Index task (v4.3.2 model pages, read Sep 28), under the $3.26 and $7.63 those pages show for Astra and Fable 5.1 at max. Change the effort level and the ranking changes with it.
It is per attempted task. AA’s methodology prices the tokens a model consumed and divides “by task count”, pass or fail. “Because cost reflects actual token usage, models that produce longer answers or more reasoning tokens will have a higher cost per task, even at identical per-token prices.” Every AA figure in this piece is AA’s cost per Intelligence Index task, or per Coding Agent task, per attempt. The method below is stricter: it divides by tasks that finished.
It drifts. v4.3 swapped in Terminal-Bench 4.0 and AutomationBench-AA, and the index has since moved to v4.3.2. We don’t compare per-task costs from before v4.3 with those after it; that’s our inference from AA’s changelog, not AA’s claim. Published figures move too. AA’s Sep 9 Astra benchmark said Astra (max) “costs $7.09 per task” on the Coding Agent Index, and the index page showed $7.47 on Sep 28. Quote an article with its date or a page with its read date, never one dressed as the other.
The Coding Agent Index v1.5 is where the price sheet inverts. Read Sep 28, at max effort: Claude Code with Opus 5.5, $13.0 per task; Claude Code with Fable 5.1 (max, with fallback), $12.4; Codex with Astra, $7.47; Codex with GPT-6 Sol, $2.99. In AA’s chart data, Opus 5.5 at max used about 333k output tokens per task against about 134k for Fable 5.1. AA is plain about scope: the figure “is intended to show pay-per-token API cost, not consumer plan pricing or the full operational cost of deploying the system in production.”
Screenshot: Artificial Analysis, “AI Coding Agent Benchmarks & Leaderboard” (undated page), captured Sep 28, 2026.
Practitioners are running the same check on vendor promises. Theo’s Sep 22 video, “Elon promised this one would be good…”, carries the description “Elon claimed Grok 4.7 would be more token-efficient, and it’s less by 30 to 80%, with very few forgiving qualities”. That is Theo’s test, not ours, and the description names no task set or effort level. The shape of the complaint is the point: a promise about tokens is a claim about cost per task, and only a per-task measurement settles it.
Why the per-token price keeps losing to the per-task bill
Anthropic said it plainly on Sep 25, in Addy Osmani’s “What a task costs on Opus 5.5”: “Two models with the same cost per token can cost very different amounts on the same task.” Turns multiply input, and “On Opus 5.5, an output token costs 100 times a cache read.” In the post’s list-price illustration, a 40-turn task sends about 2.8M input tokens, which costs $11.20 uncached, $1.62 at a 90% cache hit rate and about $0.99 at 96%. Then the line to tape above the desk: “A retry costs more than those savings.”
Screenshot: claude.dev Blog, “What a task costs on Opus 5.5” (Sep 25, 2026), captured Sep 28, 2026.
The same post shows why every per-task claim needs its setting printed beside it: “You may have seen that Opus 5.5 costs 40% less to run than Opus 5. That’s our estimate for typical workloads billed by token, at default settings.” At max effort, AA’s Sep 22 article calls Opus 5.5 “Level with Opus 5 on cost per task” ($5.98 against $5.86 per Intelligence Index task on the model pages, v4.3.2, read Sep 28). At each model’s default, those pages have Opus 5 (high) at $3.61 and Opus 5.5 (medium) at $1.34, 63% lower by our arithmetic. Three numbers, three settings, all true.
Same token price, same score, 57% apart (max vs max, Sep 7). The third tile has a different date and basis, and says so.
Step 1: Freeze a task set you can re-run
The unit of comparison is your work, not an index. Pull 20 to 40 tasks from last month’s real queue (a sizing suggestion, not a rule), weighted by what the fleet actually does: bug fixes that arrive with a failing test, small features, refactors, a slice of a migration, the overnight batch jobs. Anthropic’s post suggests running the same real task on both models, “three or four tasks before you draw a conclusion”. Treat that as the floor for a quick check. A lane decision needs more.
Each task gets a directory holding five things:
- the starting commit SHA and a clean-checkout script
- the prompt, word for word, as the lane would receive it
- the finish check (Step 2)
- an attempt budget and a wall-clock budget
- the task class (PR, overnight job, research), so results roll up later
Version the set (tasksets/2026-09.v1/) and treat any edited task as a new version. Results from v1 and v2 don’t compare, for the same reason we don’t compare costs across AA’s index versions.
Step 2: Write the finish line before the first run
“Completed” needs a definition a script can evaluate, written before anyone sees a result. For a PR task, completed means the finish check passes on the final commit and a human merges it. For an overnight job, it means the output passes validation by morning with no human rescue. Everything else is an attempt: a run that stops early, a run that ends in a plan and “next steps for you”, a run that passes by editing the test.
Write the check as a command with an exit code: npm test exits 0, the migration script reports zero pending rows, the diff touches only allowed paths. A narrated claim of success is not a finish.
The finish check is a guardrail, and guardrails fail. A check that passes on an empty diff, or one the agent can edit, will certify junk at a great price per completed task. Keep the check outside the agent’s writable paths, and have a human read a sample of “completed” diffs on every run. When the sample finds junk, that task flips to failed and the run’s numbers get recomputed.
Step 3: Count every meter on every attempt
The ledger gets one row per attempt, including the attempts you’d rather forget. These are the worksheet columns.
| Column | What goes in | Where it comes from |
|---|---|---|
taskset, task_id, task_class |
Which frozen task, which version | Task set directory |
config |
The lane recipe under test, such as Sol at medium escalating to Astra | Runner |
attempt, kind |
1, 2, 3; first try, retry or escalation | Runner |
served_model, effort_served |
What actually answered, at what level | API response and run logs, never config |
input_tokens, output_tokens |
Uncached input; output including reasoning | Usage block per request |
cache_write_tokens, cache_read_tokens |
Priced separately from input | Usage block per request |
long_context_requests |
Requests over a vendor’s context threshold | Request sizes |
tool_usd, sandbox_usd |
Web search, code execution, containers, session-hours | Vendor meters |
usd |
The attempt, priced from a dated rate file | Computed |
finish_check, pr_merged |
Pass or fail with exit code; merged or not | Check command, git host |
Price cache traffic; it isn’t free. On Anthropic’s pricing page, Opus 5.5 bills $4 base input, $5 for a 5-minute cache write, $8 for a 1-hour write, $0.20 for a cache hit and $20 output per MTok; Fable 5.1 is $10, $12.50, $20, $0.25 and $50. Tools and runtime are meters too: web search at $10 per 1,000 searches, code execution at $0.05 per container-hour past 1,550 free hours a month, Managed Agents runtime at $0.08 per session-hour. Three meters for one agent job covers the inventory; here each meter just needs a column.
One per-task trap deserves its own column. OpenAI’s GPT-6 Astra model page says “Prompts with more than 272K input tokens are priced at 2x input and cache rates and 1.5x output for the full request.” A long agent task that crosses 272K on its late turns pays the higher rate on each of those whole requests, so price per request, never from a per-task average. Astra’s short-context rates are $10 input, $1.00 cached, $12.50 cache writes and $50 output per MTok.
Where the clock and the cache move the bill, the live runbooks already cover it: off-peak scheduling by the clock, DeepSeek V4.1-Flash’s off-peak and cache-hit windows and Fable 5.1’s cache reads per overnight job. Their savings land in these columns. Speed tiers are a meter as well, and the fast-mode speed budget prices them.
Step 4: Record the effort level each run was served at
Effort moves AI agent cost per task as much as the choice of model does. AA’s per-effort pages for Opus 5.5 run from $0.55 per Intelligence Index task at low to $5.98 at max (v4.3.2, read Sep 28). The Coding Agent figures above are all max, which is rarely what a lane runs day to day. A cost row without its effort level is a number without units.
Take effort_served from run logs, not from a settings file. A session-level choice can override the file you read, and defaults move between releases, so the configured level and the served level drift apart without anyone deciding they should. Floors, ceilings and the per-model ladder belong to the sibling runbook that prices effort per completed task. This step only refuses to store a row without the level.
Step 5: Divide total spend by finished tasks
The arithmetic is one line. Sum every attempt’s dollars across the task set, including retries, escalations, cache, tools and sandboxes, then divide by the number of tasks that finished. Failed attempts stay in the numerator and leave the denominator. That is the whole difference from AA’s metric, and it’s the difference that catches a cheap model which fails half the time.
Five meters in, one division out. Failed attempts are paid for and never counted as done.
Report it two ways, dollars per merged PR for PR tasks and dollars per finished overnight job for batch work, with dollars per task tried (AA’s shape) and the completion rate beside them. The gap between the two dollar figures is your failure tax.
The script is illustrative. It assumes one JSON line per attempt with the Step 3 columns already priced.
# cost_per_task.py (illustrative): one JSON line per attempt in, one row per recipe and class out
import json, sys
from collections import defaultdict
spend = defaultdict(float) # (config, task_class) -> USD, every attempt
tried = defaultdict(set) # (config, task_class) -> task_ids attempted
finished = defaultdict(set) # (config, task_class) -> task_ids that count as done
for line in open(sys.argv[1]): # runs-2026-09.v1.jsonl
a = json.loads(line)
key = (a["config"], a["task_class"]) # config such as "sol@medium>astra@low"
spend[key] += a["usd"] + a["tool_usd"] + a["sandbox_usd"] # pass or fail
tried[key].add(a["task_id"])
done = a["finish_check"] == "pass" and (a["pr_merged"] or a["task_class"] != "pr")
if done:
finished[key].add(a["task_id"])
for key in sorted(spend):
n_tried, n_done = len(tried[key]), len(finished[key])
per_done = f"${spend[key] / n_done:.2f}" if n_done else "nothing finished"
print(*key, f"{n_done}/{n_tried} finished",
f"${spend[key] / n_tried:.2f} per task tried",
f"{per_done} per completed task")
The pr rows are dollars per merged PR; the overnight rows are dollars per finished overnight job. A recipe that finished nothing prints no price at all, which is the correct price.
Step 6: Compare models on AI agent cost per task, never on list price
The price column can’t rank models, and AA’s own pages show it twice. On the Intelligence Index, Gemini 3.8 Flash (high) is 62.5% cheaper per token than GPT-6 Sol (max) and about 17% dearer per task, by our arithmetic. On the Coding Agent Index, the order by token price and the order by task price disagree at the top.
| Model (effort) | List price, in / out per MTok | AA cost per attempted task | AA basis |
|---|---|---|---|
| Gemini 3.8 Flash (high) | $0.75 / $3.75 | $1.24 (score 41) | Intelligence Index v4.3.2, read Sep 28 |
| GPT-6 Sol (max) | $2 / $10 | $1.06 (score 48) | Intelligence Index v4.3.2, read Sep 28 |
| Claude Code, Opus 5.5 (max) | $4 / $20 | $13.0 | Coding Agent Index v1.5, read Sep 28 |
| Claude Code, Fable 5.1 (max, with fallback) | $10 / $50 | $12.4 | Coding Agent Index v1.5, read Sep 28 |
| Codex, GPT-6 Astra (max) | $10 / $50 | $7.47 | Coding Agent Index v1.5, read Sep 28 |
| Codex, GPT-6 Sol (max) | $2 / $10 | $2.99 | Coding Agent Index v1.5, read Sep 28 |
Max effort, per attempted task, read Sep 28. The cheapest token among the top three bought the dearest task.
Then pick with a rule, written before the results arrive:
- Drop any recipe whose completion rate on the task set falls below the lane’s floor. A cheap failure still costs someone the morning.
- Among the rest, the lowest dollars per completed task wins, at the effort you’ll actually run.
- If two recipes land inside your run-to-run noise, take the higher completion rate, then the shorter wall-clock where a human waits.
- Record the winner with the task set version, the rate file date and every recipe’s effort.
A cheap-model router with escalation is just another recipe in this table. The Jev routing piece already judges its tiers by cost per merged PR; this ledger is where that number comes from.
Step 7: Re-run when a model, a price or a default moves
The number goes stale on the vendor’s schedule, not yours. Re-run the task set, changing only the recipe under test, whenever:
- a new model or model version enters a lane
- a list price changes (the self-hosting runbook starts from the Sep 22 GPT-6 Sol and Luna cuts)
- a default model or default effort changes in any harness you run
- the harness or CLI version changes
- cache pricing or a long-context rule changes
- the task set gets a new version
Stamp every result with the task set version, the rate file date, the harness version, the served model and effort, and the run date. A number without those stamps is a rumor about your fleet. Run the set monthly even when nothing changed. If the number moves with no trigger, the drift is in your harness, your prompts or the vendor’s serving, and the ledger tells you which.
AI agent cost per task numbers that lie, and the signal each one leaves
| Failure | Signal | Fix |
|---|---|---|
| Dividing by tasks tried, not tasks finished | The per-task figure holds steady while the completion rate drops | Divide by finished tasks; print both figures |
| Unversioned task set | The number moves with no model, price or harness change | Freeze and version the set; an edited task is a new version |
| Effort read from config | Cost per task jumps after an upgrade with no settings change | Log effort_served from the run itself |
| Cache treated as free, or writes ignored | The ledger total lands under the invoice | Price writes and reads from the rate file |
| Long-context requests priced at the short rate | A few long tasks dominate the invoice but not the ledger | Flag requests over GPT-6 Astra’s 272K threshold and price them per request |
| Article and page figures mixed | Two numbers for one configuration in the same deck | Quote each with its own publication or read date |
| Timeouts dropped from the ledger | Attempt count below tasks times runs | A killed run is an attempt; it stays in the numerator |
| A finish check the agent can game | High completion rate, bad diffs in the human sample | Move the check out of writable paths and recompute |
The first row is AA’s metric quoted as yours; AA says exactly what it measures. The last row is how a weak model buys a good number, which is why Step 2’s human sample is part of the method rather than a courtesy.
The ledger fails quietly too: a runner that crashes mid-set writes a partial file that flatters whichever recipe ran first. Make the script refuse to print when any task and recipe pair is missing attempts, and treat a refused report as a failed run.
Cost per completed task is a fleet number, not a vendor’s
No vendor dashboard can produce this figure for you. Each one sees its own meter, prices its own tokens and counts its own attempts, while your task set runs across several harnesses, an escalation path that crosses vendors and tool meters that bill somewhere else. The ledger that joins them lives in the operating layer: the runner that replays the task set, the rate file, the finish checks, the monthly re-run. Keeping that layer is the job of a multi-agent command center, whichever vendor answered the call.
List price is one input to the number you actually pay. Treat it like one.
FAQ
What is cost per completed task for an AI agent?
It is everything a task set cost, every attempt, retry, escalation, cache write and read, and tool or sandbox meter, divided by the tasks that passed a finish check you wrote in advance. Failed runs stay in the cost and never count as done. Compare models on that figure, measured on your own tasks.
Is Artificial Analysis cost per task the same as cost per completed task?
No. Artificial Analysis divides a model’s token spend on its index by the number of tasks it ran, pass or fail, priced at API rates. That makes it a cost per attempted task. It is a useful first read on token efficiency, but only your own task set, divided by finished tasks, prices your lanes.
Sources
- Artificial Analysis: Announcing the Intelligence Index v4.3: Sep 7, 2026; Astra and Fable 5.1 both at 53, $3.26 vs $7.63 per task (max vs max)
- Artificial Analysis: methodology: cost per task divides token spend by task count
- Artificial Analysis: Coding Agent Index v1.5: API cost per task by agent, read Sep 28, 2026
- Artificial Analysis: Claude Opus 5.5 takes the top spot: Sep 22, 2026; 58 at max, “Level with Opus 5 on cost per task”
- Artificial Analysis model pages: cost per Intelligence Index task by model and effort, v4.3.2, read Sep 28, 2026 (Opus 5.5 linked; Opus 5, Gemini 3.8 Flash and GPT-6 Sol read the same day)
- Artificial Analysis: Benchmarking GPT-6 Astra: Sep 9, 2026; $7.09 per Coding Agent task at publication
- Anthropic: What a task costs on Opus 5.5: Addy Osmani, Sep 25, 2026; retries, cache and the 40% estimate
- Anthropic: Claude pricing: base input, cache writes, cache hits, output, tool and runtime meters
- OpenAI: GPT-6 Astra model page: list prices and the 272K long-context rule
- Theo: “Elon promised this one would be good…”: YouTube, Sep 22, 2026; the Grok 4.7 token-efficiency test is his
