The Evidence That Earns a Model the Default Slot

A vendor-set default model is a proposal. Pin the incumbent, then promote through shadow and canary lanes on merged-PR, revert and cost-per-task gates.

Default model promotion hero: a three-rung ladder labeled shadow, canary and default, with a pin holding the incumbent model at the foot of the ladderDefault model promotion hero: a three-rung ladder labeled shadow, canary and default, with a pin holding the incumbent model at the foot of the ladder
The vendor proposes a model. The ladder decides whether it keeps the slot.

At 15:44 UTC on Sep 22, npm published Claude Code 2.1.280, and two defaults moved with it. The opus alias now meant Opus 5.5. On Pro and Team Standard plans, the default model went from Sonnet to Opus, a tier up rather than a point release. Nobody on those plans filed a change request; the next session on an updated client simply opened on a different model.

It was the fourth time in four months that Claude Code’s default changed for at least one account type (May 28, Jul 11, Jul 24, Sep 22), and Claude Code was not even the busiest vendor this month. Copilot, Codex, Kimi and Cline each moved a default of their own in September. None of them checked your merge rate first.

Treat every vendor-set default model as a proposal. By Tuesday you want the incumbent pinned with the vendor’s own hold-back keys, five written evidence gates, and a ladder that any candidate climbs before it gets the slot: shadow, where it only logs; canary, where it writes code on a slice of lanes; then default, with a demotion trigger armed the day it arrives.

Sep 22–28: the default moved first, and the brake shipped three days later

Anthropic launched Opus 5.5 on Sep 22 at $4/$20 per million tokens, against Opus 5’s $5/$25, and says it “costs 40% less to run than Opus 5”. The launch post never mentions Claude Code’s default. The changelog does: 2.1.280 made claude-opus-5-5 “now the default Opus model” and “Changed the default model on Pro and Team Standard plans from Sonnet to Opus, matching Max, Team Premium, and Enterprise”.

Claude Code changelog entry 2.1.280 dated September 22, 2026, including the line that changed the default model on Pro and Team Standard plans from Sonnet to Opus Screenshot: Claude Code Docs, “Claude Code changelog” (2.1.280 entry, Sep 22, 2026), captured Sep 28, 2026.

Two exemptions matter for your inventory. Microsoft Foundry kept Sonnet 4.5 as its default and Opus 4.6 behind opus, per the model-config docs. And “Opus 5.5 requires Claude Code v2.1.280 or later”, while npm’s stable tag, which the setup docs describe as “a version that is typically about one week old, skipping releases with major regressions”, pointed at 2.1.274 early on Sep 28 and at 2.1.277 by that evening. Same fleet, same week, two defaults, depending on which channel a laptop tracks.

Anthropic’s head of product for Claude Code announced the flip that afternoon, with an effort default attached:

The effort moved with the model. Opus 5.5 “defaults to medium, one level below Opus 5’s default of high”, and 2.1.280 stopped applying effort levels saved before per-model effort to newly released models. Compare two defaults and you are comparing two models at two efforts.

Then the brake. Until 2.1.283 shipped on Sep 25, an allowlist entry admitted its successors: the docs say an availableModels entry such as claude-opus-5 “also permits later releases that extend it, such as Opus 5.5, as soon as Claude Code supports them.” 2.1.283 added two managed settings to hold a release back, deniedModels and availableModelsMatch set to "exact". For three days no allowlist could exclude it: a team that had carefully allowlisted Opus 5 had allowlisted Opus 5.5 too, and still has unless it sets the new key.

Claude Code model configuration docs, section Block specific models or versions, explaining that an availableModels entry such as claude-opus-5 also permits Opus 5.5 and listing the deniedModels and availableModelsMatch settings Screenshot: Claude Code Docs, “Model configuration” (undated page), captured Sep 28, 2026.

Claude Code moved again on Sep 28, when 2.1.284 added Sonnet 5.5, “now the default Sonnet model on the Anthropic API”. The rest of the stack moved on its own clocks. GitHub’s Sep 22 changelog says “new models are automatically enabled unless an administrator has turned off the global default or explicitly disables this model.” Codex CLI 0.154.0 (Sep 9) made fresh sessions “respect server model defaults unless explicitly overridden”, per the Codex changelog.

Kimi Code 0.42.0 (Sep 9) raised the default thinking effort “to the recommended level for eligible users”. Cline’s catalog refreshes changed the resolved default for 44 providers on Sep 17, 36 on Sep 22 and 19 on Sep 24; its GitHub Copilot provider moved to GPT-6 Astra, then to Claude Opus 5.5 seven days later. The v4.1.21 notes say it plainly: “expect a different default.”

Not every launch is a default change. MiniMax announced M3.1-Flash-Preview on Sep 27, “available only through Token Plan and MiniMax Code for now” per its model docs, and at least one write-up that day called it MiniMax Code’s new default. MiniMax’s own CLI says otherwise; PR #375 reads: “The default model remains MiniMax-M3, and an existing managed snapshot remains authoritative.” What MiniMax does default is effort: on its API, an omitted effort means max.

The description of Theo’s Sep 23 video says Opus 5.5 “got a full day of real work done on about 20% of my weekly Claude limit”. That is his view of his day: a fair reason to shadow a model next, and not a gate, with no task list, token counts or revert window.

Anthropic’s API guidance sides with the operator, even where its CLI did not wait. The Opus 5.5 migration guide says “Test in a development environment before switching production traffic.” The launch post itself concedes that “benchmark margins have become a less reliable guide to real-world differences.” No vendor publishes your merged-PR or revert rate, so the gates below are our practice, not vendor guidance.

Step 1: Inventory every default model and effort a vendor can move

Agents act on the default. A chat user who gets a new model notices the tone; a lane that gets one opens pull requests with it, at whatever effort the vendor chose, under your name. So the first artifact is a list, per harness and per plan, of every default the vendor controls: the model, where each alias points, the effort, and whether new models switch themselves on.

September’s changes, in one picture:

Timeline of vendor-set default model and effort changes in September 2026: Codex CLI 0.154.0 and Kimi Code 0.42.0 on Sep 9, Cline v4.1.19 on Sep 17, Claude Code 2.1.280, Copilot default enablement and Cline v4.1.20 on Sep 22, Cline v4.1.21 on Sep 24, the Claude Code 2.1.283 hold-back keys on Sep 25, MiniMax’s Sep 27 launch with its default unchanged, and Claude Code 2.1.284 making Sonnet 5.5 the default Sonnet model on Sep 28Timeline of vendor-set default model and effort changes in September 2026: Codex CLI 0.154.0 and Kimi Code 0.42.0 on Sep 9, Cline v4.1.19 on Sep 17, Claude Code 2.1.280, Copilot default enablement and Cline v4.1.20 on Sep 22, Cline v4.1.21 on Sep 24, the Claude Code 2.1.283 hold-back keys on Sep 25, MiniMax’s Sep 27 launch with its default unchanged, and Claude Code 2.1.284 making Sonnet 5.5 the default Sonnet model on Sep 28 Eight vendor-set default changes between Sep 9 and Sep 28, one hold-back control, and one launch that left the default alone.

Turn the timeline into a pin sheet, one row per harness:

Harness What the vendor changed (date) What you pin Hold-back control
Claude Code 2.1.280 (Sep 22): opus alias to Opus 5.5; Pro and Team Standard default from Sonnet to Opus; Opus 5.5 effort defaults to medium. 2.1.284 (Sep 28): Sonnet 5.5 becomes the default Sonnet model on the Anthropic API Full model ID, or ANTHROPIC_DEFAULT_OPUS_MODEL for the alias; an explicit effort Managed availableModels with availableModelsMatch: "exact", or deniedModels (2.1.283+, Sep 25)
GitHub Copilot Sep 22: new models enabled by default under default model enablement The models each lane may use An admin turns off the global default or disables the model
Codex CLI 0.154.0 (Sep 9): fresh sessions follow server model defaults; PR #47332 (Sep 22): migration prompts to GPT-6 Sol and Luna, retired GPT-5.4 selections retargeted Model and effort, set explicitly per lane Explicit override; Enterprise admins must enable new models
Kimi Code CLI 0.42.0 (Sep 9): default thinking effort raised for eligible users Effort, explicitly None named in the changelog; confirm effort in transcripts
Cline v4.1.19, v4.1.20, v4.1.21 (Sep 17, 22, 24): resolved default changed for 44, 36 and 19 unpinned providers A model on every provider you use Pinning; unpinned providers get the new default
MiniMax Code Sep 27: M3.1-Flash-Preview announced; CLI default stays MiniMax-M3 (PR #375) Model, plus effort on every API call, since omitted means max An existing managed snapshot stays authoritative
Gemini API latest aliases are hot-swapped with every new release A specific stable model Email notice two weeks ahead, for breaking changes only

Two rules fall out of that table. Every lane names its model by full ID, because the model-config docs say “Aliases point to the recommended version for your provider and update over time.” And every lane names its effort, because at least two of September’s changes moved effort along with the model or instead of it. How to pick the level is its own exercise: price effort per completed task before you write the number down.

Upgrades are a separate check: canary every CLI upgrade against its lane manifest and fail it on anything the manifest did not declare. This runbook starts where that one stops, with a new model asking for the slot.

Step 2: Pin the incumbent with the vendor’s own hold-back keys

Where you administer Claude Code through managed settings, the pin is a file. The keys below are the ones the model-config docs and the 2.1.283 changelog name; the values are illustrative, for a fleet whose incumbent is Opus 5:

{
  "requiredMinimumVersion": "2.1.283",
  "model": "claude-opus-5",
  "availableModels": ["claude-opus-5"],
  "availableModelsMatch": "exact",
  "enforceAvailableModels": true
}

With "exact", per the changelog, “an availableModels entry allows only the model version it names, so new releases stay blocked until listed.” enforceAvailableModels extends the allowlist to the Default option. deniedModels is the alternative when you would rather name the single release to block. Canary hosts get their own copy of the file, with the candidate added to availableModels; production hosts never do.

Outside managed settings, pin per lane: the full ID in --model, ANTHROPIC_MODEL or the settings model key. A lane that pins opus is pinned to whatever opus means this week.

Now the part most pins skip: what happens when the pin itself fails.

  • Old clients ignore both keys. The docs tell you to set requiredMinimumVersion as well, because versions before 2.1.283 do not read deniedModels or availableModelsMatch, and the floor keeps those versions from starting. Check what it does to laptops on the stable channel before you ship it; stable reached only 2.1.277 on Sep 28.
  • A blocked Default steps down instead of stopping. If the Default model is blocked, Claude Code falls to the newest permitted version of the same family, then Sonnet, then Haiku, then the first availableModels entry; it refuses to start only when none of those is permitted. A pin written wrong can land a lane on a different family and keep it running.
  • An Enterprise organization default is a suggestion. The docs call it “a starting point, not a restriction”, read once at startup. Only the managed keys restrict.

The wall behind the pin is a log: every call records the requested model beside the one that answered, and a mismatch pages the lane owner. That log is the served-model check from this batch; build it before you trust any pin, because it is also the first gate on the ladder.

Step 3: Write the evidence gates before the candidate runs

Write the thresholds while you still have no favourite. A threshold written after the canary is a description of the canary. Five gates, plus a sample floor, and every number below is illustrative: set yours from the incumbent’s last 30 days on the same lanes.

Gate Metric Threshold (illustrative) Source
Served model Share of calls whose response names the candidate’s pinned ID 100%; any unexplained mismatch stops the clock Served-model log: response model field, subagent transcripts
Pinned effort Effort level recorded on every run Set explicitly and identical across the rung; never the vendor default Lane config, checked against the transcript
Cost per completed task Spend divided by tasks merged and not reverted Shadow: at most 1.25x the incumbent. Canary: at most 1.10x Billing export joined to the task ledger
Merged-PR rate PRs merged divided by PRs the lane opened Canary only: no more than 5 points below the incumbent Git host API
Revert rate Merged PRs reverted or hot-fixed within 14 days Canary only: at or below the incumbent Revert commits and linked fix PRs
Sample floor Completed tasks and calendar days at the rung Shadow: 30 replayed tasks. Canary: 30 completed tasks and 14 days Task ledger

The cost gate divides by finished work, not tokens, for the reasons in cost per completed task: a cheaper model that needs two attempts is not cheaper. The effort gate exists because a default can move effort under a model you never touched, as Kimi 0.42.0 did on Sep 9.

Default model promotion ladder: a vendor-set default enters as a proposal while the incumbent stays pinned, climbs to shadow where it only logs, clears gate one into canary on a slice of lanes, clears gate two into default, and a demotion trigger hands the slot back to the pinned incumbentDefault model promotion ladder: a vendor-set default enters as a proposal while the incumbent stays pinned, climbs to shadow where it only logs, clears gate one into canary on a slice of lanes, clears gate two into default, and a demotion trigger hands the slot back to the pinned incumbent Two gates up, one trigger down. The incumbent keeps the slot until the candidate clears both.

If you already fail pull requests on eval regressions, eval regression gates in CI are the per-PR half of this discipline. The ladder is the per-model half, and it needs different evidence: merges and reverts over weeks, which no eval suite produces.

Step 4: Shadow: the candidate runs, logs and never merges

Replay last week’s completed tasks through the candidate. Same repo commit, same prompt, same tools, the effort you pinned, a scratch worktree, and no push credentials. The output is a diff, a test result, a cost line and a served-model line per task. Nothing it writes reaches a branch anyone else can see.

Score each replay against the incumbent’s merged result: did the candidate’s diff pass the tests the merged PR passed, at what cost, on which model. A task the candidate errors on counts as failed, not skipped; a shadow that drops its failures will always look good.

Shadow is where vendor claims meet your repo. The model-config docs say “In Anthropic’s testing, Opus 5.5 at medium matches or exceeds Opus 5 at high on coding and knowledge-work evaluations”, and the prompting guide says to “test several levels against your own evals”. Run the candidate at two efforts and the incumbent at its current one, and let the cost gate pick. On Claude Code, 2.1.283’s /doctor prompt-audit is worth a pass here too, since it flags prompting patterns written for older models.

The pattern is the one a shadow-mode CI gate uses for a classifier: run, log, compare, never act. Exit shadow when gate one clears: served model, effort, cost at 1.25x or better, and 30 replayed tasks.

Step 5: Canary: a slice of lanes, real merges, the same review

Give the candidate one or two lanes whose mistakes are cheap to undo: tests, docs, internal tooling. Review stays exactly as strict as on any other lane, and nothing auto-merges.

Tag every PR with the served model and effort, in a commit trailer or the PR body, so the revert query in two weeks can find them. Run the incumbent’s lanes over the same window and compare like with like. A merge only counts as good after 14 days without a revert or a hot-fix.

The rollback is one edit to the canary copy: take the candidate out of availableModels and point model back at the incumbent. Confirm where the lanes landed in the served-model log, not in the settings file; a blocked default steps down, and not always to the model you expected.

Step 6: Promote with a record, and arm the demotion trigger the same day

Promotion is a change to the production managed-settings file with a name on it: the candidate goes into availableModels and model, and the incumbent stays listed for 30 days so that demotion is also one edit. Write a promotion record with the date, the candidate’s full ID, the pinned effort, each gate’s number against the incumbent’s, and the approver.

The demotion trigger is the same gate check, run weekly against a rolling 14-day window at the default rung. Here is an illustrative version, in Python, that turns one rung’s summary into a verdict:

# gate_check.py (illustrative): one rung summary in, one verdict out
import json, sys

def verdict(rung):
    c, i = rung["candidate"], rung["incumbent"]
    gates = {
        "served_model": c["served_match_rate"] == 1.0,
        "effort_pinned": c["effort_source"] == "pinned",
        "cost_per_completed": c["usd_per_completed"] <= 1.10 * i["usd_per_completed"],
        "merged_pr_rate": c["merged_rate"] >= i["merged_rate"] - 0.05,
        "revert_rate_14d": c["revert_rate_14d"] <= i["revert_rate_14d"],
        "sample": c["completed_tasks"] >= 30 and c["days"] >= 14,
    }
    failed = [name for name, ok in gates.items() if not ok]
    if not failed:
        return "PROMOTE", failed
    if not gates["served_model"] or c["revert_rate_14d"] > 1.5 * i["revert_rate_14d"]:
        return "DEMOTE", failed
    return "HOLD", failed

try:
    print(*verdict(json.load(open(sys.argv[1]))))
except (OSError, IndexError, KeyError, TypeError, ValueError) as err:
    print("HOLD", ["missing_or_bad_data", repr(err)])  # a broken check never promotes

Three properties matter more than the thresholds. Missing data returns HOLD, so a broken ledger can stall a promotion but never cause one. A served-model mismatch or a revert spike demotes without waiting for a meeting. And the script only prints: the enforcement is the settings file and branch protection, and the human who edits the file owns the call.

How a default model slips past the ladder, and the signal for each

The version gap. A client on 2.1.280, 2.1.281 or 2.1.282 already defaults to Opus 5.5 and does not read the hold-back keys. Signal: a lane inventory with no version column, or production lanes serving the candidate while the settings file says otherwise. Fix: requiredMinimumVersion, and the version recorded beside every lane.

The step-down to a cheaper family. A pin blocks the Default model and the lane lands on Sonnet or Haiku. Signal: a different model family in the served-model log, with no settings change to explain it. Fix: set model explicitly as well as the allowlist.

Effort drift under a pinned model. The model held and the effort moved. Signal: output tokens per task jump or drop with no model change. Fix: effort in the lane config, checked against the transcript.

The unpinned default that lands on nobody’s model. Cline’s Sep 24 CLI refresh made “Space Bunny Free” the resolved default for its OpenCode Go provider, and Space Bunny is a stealth model with no named maker. Signal: a served model you cannot find on any price page. Fix: pin every provider, and send stealth names through the rumor-intake rule before they get near a ladder.

The canary that proves nothing. Easy tasks, a short window, or a lane nobody reviews closely. Signal: the sample floor met in days on a lane that usually takes weeks, or a task mix unlike the incumbent’s. Fix: the same task mix on both sides, and treat the floor as a floor.

The default slot belongs to the fleet’s evidence

Every vendor on the pin sheet moves its default on its own schedule, and none of them sees your merge rate. So the ladder cannot live inside any one vendor’s product. It lives in the layer that runs the fleet: a managed-settings file per host class, the served-model log, the task ledger joined to billing, the weekly gate check and a promotion record with a name on it. That layer is what a multi-agent command center is once you take the dashboard off it.

The next vendor default is probably already in a changelog you have not read. It gets a row on the pin sheet and a place at the foot of the ladder.

FAQ

How do I stop Claude Code from switching to a new default model?

Pin the full model ID with --model, ANTHROPIC_MODEL or the settings model key, never an alias. If you administer managed settings, ship availableModels plus availableModelsMatch: "exact", or deniedModels, and set requiredMinimumVersion, since versions before 2.1.283 ignore the hold-back keys. Then confirm it in your served-model log.

How long should a new model stay in canary before it becomes the default?

Long enough for reverts to show up. Our practice is at least 14 days and 30 completed tasks on one or two low-risk lanes, compared against the incumbent’s lanes over the same window. A merge counts as good only after 14 days without a revert or hot-fix. The numbers are illustrative.

Did MiniMax make M3.1-Flash-Preview the default in MiniMax Code?

No primary MiniMax source says so. PR #375, merged Sep 28, states that the default model remains MiniMax-M3 and that an existing managed snapshot stays authoritative. MiniMax does document a default effort: on its API, omitting effort means max, so set effort explicitly on every call.

Sources

YOU'RE THROUGH THIS ONE.

Keep connecting the dots.

Back to the library