Kill Low-Effort Retry Loops: Price Effort per Completed Task

Reasoning effort cost looks cheap per attempt until you count retries. Record served effort, price effort per completed task, then pin a floor and ceiling.

Effort per completed task: a five-rung effort ladder from low to max, with a retry arrow looping beside the bottom rungEffort per completed task: a five-rung effort ladder from low to max, with a retry arrow looping beside the bottom rung
Low is the cheapest rung per attempt. Per finished task, it depends on how often you climb it twice.

On Sep 27, eighteen worker attempts in one managed project all ran GPT-5.6 Terra at low reasoning effort. Thirteen came back blocked, two exited early and three finished. The owner’s issue asked to raise the minimum routed effort to medium for every model, and the fix merged on Sep 28, with an honest footnote: brief and environment defects also contributed, “so this is not a controlled comparison.”

Almost nobody has one, because the meters price effort per token and per attempt, never effort per completed task. Reasoning effort cost only becomes a real number when you divide it by finished work, with every retry, blocked run and early exit in the numerator. Do that and low effort stops looking cheap on some lanes while staying cheap on others, and you can tell which is which.

This is the runbook for Tuesday. For each lane you record the effort the model actually served, run a fixed task set at two or three levels, compute tokens and dollars per completed task, set a floor and a ceiling, and pin the level per model. Then you re-run it whenever a vendor moves a default, which in September happened about once a week.

Sep 22: Opus 5.5 ships at medium effort, and the defaults stop agreeing

Claude Code v2.1.280 shipped on Sep 22 with Claude Opus 5.5 as the default Opus model, and it moved the default on Pro and Team Standard plans from Sonnet to Opus. Opus 5.5 runs at medium unless something sets a level for it. The model configuration docs say “In Anthropic’s testing, Opus 5.5 at medium matches or exceeds Opus 5 at high on coding and knowledge-work evaluations,” without publishing the numbers. Most other models that support effort still default to high.

The settings rule is where fleets get caught. Per the docs, “a top-level effortLevel in your user settings file doesn’t count for Opus 5.5”, so an old high in that file no longer reaches the model you now run by default. The same key behaves the opposite way one file over: “A top-level effortLevel in project, local, or managed settings, or one passed with --settings, applies to every model.”

Claude Code Docs model configuration page, open at the Available models section, with the settings sidebar on the left and an on-page contents list covering model aliases, model restrictions and the organization default model Screenshot: Claude Code Docs, “Model configuration - Claude Code Docs” (undated page), captured Sep 28, 2026.

On Sep 28, npm’s stable tag for Claude Code moved from 2.1.274 to 2.1.277, and Opus 5.5 requires 2.1.280 or later, so stable-channel lanes have not met the new default yet. They will, on a day nobody scheduled.

Codex says it differently. OpenAI’s model guidance reads “Start with Medium for Sol, High for Luna, or Light for Astra,” where Light is low in configuration, while the app’s starting preset is Sol Light. It adds that “Most tasks do not need Max or Ultra,” and that efforts don’t map exactly between model generations. The Codex lead made the point on X on Sep 6:

Anthropic’s two effort posts landed on Sep 25. In “What a task costs on Opus 5.5”, Addy Osmani prices the trade in two sentences: “So high pays for itself on a task where it saves one retry. On a task medium would have finished the first time, it’s wasted.” In “Spending your effort”, Thariq Shihipar reports internal Terminal-Bench 3.0 runs in which Fable 5.1 passed 140 of 370 attempts at low and 214 of 370 at max, and notes that more effort mostly fixed missed edge cases, not a wrong approach.

A paper had already questioned the premise that low is cheap. ScrambleToolBench, submitted Aug 3, puts it flatly: “Lower reasoning does not reliably reduce token cost.” For Claude Sonnet 5 under its mapping-drift condition, low used 13,274 output tokens per solved task, against 9,741 at medium and 11,652 at high.

Why a cheap attempt is not a cheap agent run

In a chat window, a weak answer costs you a reread. In an agent lane it becomes a tool call, a test run, a blocked state, a retry, and eventually a person reading a transcript. ScrambleToolBench shows the mechanism: at low effort, 7 of 19 retained drift episodes ran into the 100-action limit, against 1 of 20 at medium and none at high. Low effort saved on thinking and spent the savings on actions.

Bar chart of output tokens per solved task for Claude Sonnet 5 under mapping drift in ScrambleToolBench: 13,274 at low, 9,741 at medium and 11,652 at high reasoningBar chart of output tokens per solved task for Claude Sonnet 5 under mapping drift in ScrambleToolBench: 13,274 at low, 9,741 at medium and 11,652 at high reasoning ScrambleToolBench, arXiv 2608.02358 (Aug 3, 2026). Claude Sonnet 5, mapping-drift condition only; about 20 five-task episodes per cell, failed-episode tokens included.

That is why the vendor advice for unattended work is a stop rule. Anthropic’s Opus 5.5 prompting guide tells loops to “stop after two or three automatic continuations on the same task rather than repeating them indefinitely,” and a low-effort retry loop is the most expensive way to reach that stop.

The loop that prices effort per completed task

Effort per completed task pricing loop: record served effort, run the fixed task set at two or three levels, compute dollars and tokens per completed task, set floor and ceiling, pin per model, and re-run on any default changeEffort per completed task pricing loop: record served effort, run the fixed task set at two or three levels, compute dollars and tokens per completed task, set floor and ceiling, pin per model, and re-run on any default change The loop closes on purpose. A vendor default change is a new measurement, not a footnote.

Step 1: Record the effort each lane actually served

Config is a request. The served level is what ran. Claude Code resolves effort in a fixed order: an explicit choice (CLAUDE_CODE_EFFORT_LEVEL, --effort or /effort), then settings (the per-model saved level or effortLevel), then the model’s default. Two quieter overrides sit on top. Skill and subagent frontmatter effort overrides the session level, and an organization cap from maxEffortLevel clamps whatever was asked for; per the docs, with json or stream-json output or in background agents, “the clamp applies silently”.

So the settings file is the wrong source of truth. Have the launch wrapper set the level explicitly on every run and log it with the model, the CLI version and any cap in force. An explicit choice outranks the settings files and the default; skill or subagent frontmatter and the cap are the exceptions, so log those too. The rows below are illustrative; copy the columns.

Lane Model Served effort (logged) Source of truth What can change it silently
Desk pairing Opus 5.5, Claude Code medium wrapper sets CLAUDE_CODE_EFFORT_LEVEL; run log an org maxEffortLevel cap
CI fix loop Opus 5.5, Claude Code headless high --effort in the launch line; run log the cap, silently, in stream-json
Test-writing subagents Opus 5.5 via subagent file low frontmatter effort, checked into the repo a teammate edits the file
Review lane GPT-6 Astra, Codex low level the wrapper passes; run log a launch from the app, which starts at Sol Light
API batch workers Opus 5.5, Claude API medium effort field in your request log a code path that omits the field

Once a week, diff the logged level against what the settings files claim. The user-settings effortLevel that Opus 5.5 ignores is the textbook mismatch: the file says high, the model runs medium, and every cost number built on the file is wrong. “Unknown” is a valid entry; it means unmeasured, not fine.

Step 2: Freeze a task set with finish lines a script can check

Pick eight to twelve real tasks per lane class from last month’s history. Each needs a finish line a script can check without a human: tests green, a PR opened with CI passing, a diff that applies cleanly. Mix easy and long-horizon work in the proportion the lane actually sees, because a set of easy tasks will always crown low effort.

Freeze everything except effort: repo commit, prompt, tool list, model version and CLI version. Headless lane reproducibility covers the pinning; a run that drifts mid-comparison measures the drift. Anthropic’s Sep 25 cost post asks for “three or four tasks before you draw a conclusion” and suggests a minimal ladder: “Run one hard task at medium and then at high. Run one mechanical task at low.” That is the smoke test. The frozen set is the measurement.

Step 3: Run two or three levels and price effort per completed task

Run the set at the lane’s current level and one level either side. Keep every attempt, including the blocked, the abandoned and the ones a human killed. Then compute, per model and level: attempts, completions, dollars per completed task (all dollars ÷ completions), output tokens per completed task, and attempts per completion. Anything above two attempts per completion is a retry loop wearing a budget label.

Vendor numbers only cover the per-attempt half. Artificial Analysis prices Opus 5.5 at each effort level on its model pages (v4.3.2, read Sep 28), and those figures are its cost per Intelligence Index task, computed from token spend “then dividing by task count” per its methodology. Read the slope rather than the points: low to medium buys 9 index points for 79 cents, while xhigh to max buys 2 points for $2.52 (our arithmetic).

Chart of reasoning effort cost for Claude Opus 5.5 on the Artificial Analysis Intelligence Index v4.3.2: score against dollars per attempted task at low, medium, high, xhigh and max effortChart of reasoning effort cost for Claude Opus 5.5 on the Artificial Analysis Intelligence Index v4.3.2: score against dollars per attempted task at low, medium, high, xhigh and max effort Artificial Analysis, Intelligence Index v4.3.2, read Sep 28, 2026. Per attempted task, pass or fail. Not cost per completed task.

AA ran all five levels with Anthropic’s default fallback enabled. Its frontier ranks models against each other, not your lane.

Artificial Analysis Opus 5.5 write-up listing key takeaways, including that max, xhigh, high and medium sit on the intelligence versus cost per task frontier and that all five effort settings ran with default fallback enabled Screenshot: Artificial Analysis, “Claude Opus 5.5 takes the top spot on the Artificial Analysis Intelligence Index” (Sep 22, 2026), captured Sep 28, 2026.

Your run supplies the other half. This worksheet is illustrative, with round numbers for one long-horizon lane: ten tasks, up to three attempts each at the same level.

Level $ per attempt First-try success Attempts Completed Spend $ per completed task
low $0.60 30% 21.9 6.6 $13.14 $2.00
medium $1.20 70% 13.9 9.7 $16.68 $1.71
high $1.70 80% 12.4 9.9 $21.08 $2.13

Medium wins, and low loses twice: it costs more per finished task and leaves more than three tasks for a person to finish. Run the same arithmetic on a mechanical lane (illustrative: low at $0.20 per attempt and 90% first-try success, medium at $0.45 and 95%) and low costs less than half as much per completed task. Same model, opposite answer. That is why the floor is per lane.

This illustrative script produces those columns from one JSON line per attempt:

# effort_price.py (illustrative): one JSON line per attempt in runs.jsonl
# {"lane":"ci-fix","task":"t07","model":"claude-opus-5-5","effort":"medium",
#  "attempt":2,"usd":1.12,"output_tokens":41250,"completed":true}
import json
from collections import defaultdict

agg = defaultdict(lambda: {"usd": 0.0, "out": 0, "attempts": 0, "tasks": set(), "done": set()})
for line in open("runs.jsonl"):
    r = json.loads(line)
    a = agg[(r["lane"], r["model"], r["effort"])]
    a["usd"] += r["usd"]
    a["out"] += r["output_tokens"]
    a["attempts"] += 1
    a["tasks"].add(r["task"])
    if r["completed"]:
        a["done"].add(r["task"])

for (lane, model, effort), a in sorted(agg.items()):
    done = len(a["done"])
    if not done:
        print(f"{lane} {model} {effort}: 0/{len(a['tasks'])} completed, ${a['usd']:.2f} spent")
        continue
    print(f"{lane} {model} {effort}: {done}/{len(a['tasks'])} completed, "
          f"{a['attempts'] / done:.1f} attempts per completion, "
          f"${a['usd'] / done:.2f} and {a['out'] // done:,} output tokens per completed task")

Step 4: Set a floor and a ceiling for each lane class

A floor is the lowest level a lane may run; a ceiling is the highest it may reach without a written reason. Start from the defaults below and replace each cell with what step 3 measured. The spine is Thariq Shihipar’s rule of thumb from Sep 25: low for in-the-loop sketching, medium for most feature work, high where verification matters, max for fully autonomous hard problems. The unattended floors are ours; the Sep 27 report and ScrambleToolBench both point there.

Lane class Floor Ceiling Evidence that moves it
Mechanical: renames, lint fixes, lookups, subagent chores low medium low’s cost per completed task stays under medium’s
Interactive feature work, a person in the loop medium high the person stops correcting at the higher level
Verification-heavy: migrations, security fixes, test suites high xhigh fewer reopened PRs, not just more tokens
Long-horizon unattended: overnight, multi-hour medium; never low if completions drop xhigh completions per dollar on the frozen set
Hard autonomous problems high max, one session, with a named reason tasks max finishes that xhigh does not

The ceiling is where opinions collide. AA’s index has max ahead of xhigh (58 against 56), while Theo reported the opposite in his own testing:

That is his measurement on his tasks, and AA’s is theirs on theirs; two credible rankings in opposite order argue for step 3, not for a side. In Claude Code the ceiling’s wall is maxEffortLevel in managed settings, which clamps silently in headless output. And max applies to the current session only unless it arrives through CLAUDE_CODE_EFFORT_LEVEL, so that variable is the path to a standing max, and the one to grep for in CI configs.

A floor fails open: launch a lane by hand, skip the wrapper, and the model default applies (medium on Opus 5.5, high on most others). The walls behind it are the retry cap in step 6 and the per-task budget.

Step 5: Pin effort per model, not per fleet

Pin levels per model, because the fleet-wide key now means two different things. A project-level effortLevel: low written for a cheap model also drags Opus 5.5 down, since outside the user settings file the top-level key “applies to every model.” The per-model form is modelSettings. This is the documented example from Claude Code’s settings reference, which keeps Opus 5.5 at high while other models use their own saved or default levels:

{
  "modelSettings": {
    "claude-opus-5-5": {
      "effortLevel": "high"
    }
  }
}

max isn’t accepted in modelSettings or in effortLevel, and an explicit --effort or environment variable still wins over the file.

Other harnesses need the same pin in their own dialect. In Codex, set the level per model in the wrapper and re-measure when a lane changes model. On the Claude API, send effort on every request and keep it fixed per lane, because the prompting guide warns that changing the top-level value between requests invalidates the prompt cache. Claude Code on an API key or subscription keeps the cache through an effort change; on Bedrock, Google Cloud or a gateway it “still clears the cached conversation”, per the Sep 25 cost post. On overnight lanes that live on cache reads, that is a line item.

Step 6: Kill retry loops with one escalation, not a third attempt

Allow one automatic retry at the same level. If the second attempt fails, make one attempt at the next level up. If that fails too, stop and hand the transcript to a person. Never run a third attempt at low.

The arithmetic favors escalation. Anthropic’s illustration, built from list prices, puts 20K extra thinking tokens at high at $0.40, about the same as a ten-turn retry loop at 100K cached context. One step up costs about one loop, and it pays for itself if it saves a single retry. This illustrative policy file is read by the launch wrapper, never by the model:

# lane-effort.yaml (illustrative)
lane: ci-fix
model: claude-opus-5-5
floor: medium
ceiling: xhigh
start_at: medium
retry:
  same_level_attempts: 2 # the first try plus one retry
  escalate_to: high      # exactly one attempt, one level up
  then: stop_and_page    # a person reads the transcript
budget_per_task_usd: 6.00
timeout_minutes: 45

The wrapper counts only the retries it launches. A harness that auto-continues internally slips past that count, so the budget and the timeout are the wall behind the rule. CI fix-loop guards covers the same wall for pipelines.

Step 7: Re-run the loop on every default change

Re-run the frozen set whenever a default can move under you: a CLI release that touches models or effort, the stable channel catching up, a new model, a new organization cap, or a vendor changing its starting presets. Run it at the lane’s current level and one either side, compare cost per completed task with last time, and move the pin only when the answer moved. CLI upgrade canaries catch the drift itself; this step prices it.

A new model earning the default slot needs more than an effort re-run, and the promotion evidence is its own checklist. Effort is also one input among several: the pillar on cost per completed task covers the rest, and speed has its own meter in the fast-mode speed budget.

Where effort per completed task goes wrong, and the signal for each

The ghost setting. A user-settings effortLevel that Opus 5.5 ignores. Signal: the step 1 diff shows a logged level that differs from the file. Fix: pin the model under modelSettings and delete the stale key.

The fleet-wide low. A project, local or managed effortLevel written for one model lands on every model. Signal: Opus 5.5’s completion rate drops in one repo and nowhere else. Fix: move the level into a per-model entry.

The silent clamp. An organization maxEffortLevel lowers headless and background runs without a message. Signal: a headless lane costs more per completed task than its interactive twin at the “same” level. Fix: log the cap per run, then lift it for that lane or accept the cost.

The per-attempt number in a per-completed slot. A planning doc quotes $1.34 as Opus 5.5’s cost per completed task. Signal: a dollar figure without its effort level, index version and date. Fix: send it back; AA’s numbers are per attempted task.

The stable-channel ambush. Stable installs meet Opus 5.5 and its medium default about a week late. Signal: cost per completed task moves with no config change. Fix: check the CLI version first, then run step 7.

Effort is a fleet setting, so it lives in the fleet layer

Nothing in this runbook is a prompt. The served level, the floor, the ceiling, the retry cap and the re-run trigger all live outside the model, in the wrapper that launches lanes and the log that records what ran. That is the layer a multi-agent command center is once you strip the dashboard off it: one place that knows which lane ran which model at which level, across vendors whose defaults disagree and move weekly.

Vendors tune defaults for their median user, and your lanes are not the median. Price the finished task, pin the level per model, and let the next default change start a measurement instead of an argument.

FAQ

Is high reasoning effort worth the extra cost?

Only on tasks where it prevents a retry. Anthropic’s own illustration prices 20K extra thinking tokens at high at about the same as one ten-turn retry loop. Measure cost per completed task on your frozen task set at two levels; if high cuts retries enough, it is cheaper per finished task.

Why does Opus 5.5 ignore my effortLevel setting?

Because a top-level effortLevel in your user settings file doesn’t count for Opus 5.5, which starts at medium until you choose a level for it. The same key in project, local or managed settings, or passed with --settings, applies to every model. Pin Opus 5.5 under modelSettings instead.

Does lower reasoning effort save tokens?

Not reliably. ScrambleToolBench found Claude Sonnet 5, under its mapping-drift condition, used 13,274 output tokens per solved task at low, against 9,741 at medium and 11,652 at high. Low effort hit the action limit far more often. Cheaper per step can still mean dearer per completed task.

Sources

YOU'RE THROUGH THIS ONE.

Keep connecting the dots.

Back to the library