Who Actually Answered? Log the Served Model When Safety Fallbacks Swap It
Safety fallbacks can hand your request to an older model mid-run. Log the requested vs served model on every call, alert on mismatches, and fail closed.
Go deeper. Build your own.
At 00:04 UTC on Sep 26, a Claude Code subagent that asked for opus stopped getting Opus 5.5. The user who reported it the next day counted 341 calls on claude-opus-5-5 before the switch and 321 on claude-opus-4-8 after it, while the parent session’s own report kept crediting “Opus” for the whole run. The switch surfaced only because the reporter read the subagent’s transcript one line at a time.
Nothing in that report is exotic. Safety fallbacks are a documented feature now: when a classifier flags a request, the model you asked for can hand the turn to an older one, and Anthropic’s apps label the switch. The gap is everywhere the label never reaches: subagents, headless runs, API responses nobody parses, and one desktop app where a user says nothing recorded the swap at all. If your change control has to say which model wrote a line, “the one we requested” stopped being a safe answer.
This is the runbook for closing that gap. By Tuesday, every call in your fleet logs the requested model beside the served one, taken from the response or from the subagent’s own transcript and never from a parent’s summary. A mismatch pages someone, and commits from mismatched turns carry a tag and get a second review. Lanes whose change control names the model fail closed, and one fixture proves, on a schedule, that the log still catches a fallback.
Sep 10–27: documented safety fallbacks, and two silent ones
Anthropic’s help center page on model switching, updated Sep 24, is the per-category reference. It says classifiers run “on every user request” and on what the model reads along the way: memory, connector content, web results and files. On cybersecurity it is specific: “Opus 5 or Opus 5.5 may fall back to Opus 4.8 when our cyber classifiers flag potentially higher-risk offensive cybersecurity requests”. Its examples are exploit generation, binary-based vulnerability scanning and penetration testing; secure coding and source-code vulnerability scanning stay on the model you picked.
The other categories split by model. On Opus 5.5, dual-use biology requests and frontier-LLM-development requests fall back to Opus 5. Distillation behaves differently: “Distillation blocks don’t fall back to another model, and the request is blocked outright.”
The Opus 5.5 launch page (Sep 22) disagrees on that last point: it groups the cybersecurity, biology and distillation safeguards together, “all of which fall back to another model transparently”. Use the help center for per-category behavior. Keep one word from the launch page, though: “most cybersecurity tasks will be re-routed to Opus 4.8.” Most, not all, and no share is published.
Where the label exists, it works. “All fallbacks are transparent, meaning you’ll see a notice explaining that the model switched, and the response will be labeled with the model that answered.” Then comes the sentence that matters for overnight runs: “After the switch, the model picker stays on the less capable model for the rest of the conversation.” On the API nothing switches unless you ask for it: “Automatic switching isn’t active by default, and API customers must opt into and configure the fallbacks.”
Screenshot: Claude Support, “Why Claude switched models in your conversation with Opus 5 or Opus 5.5” (updated Sep 24, 2026), captured Sep 28, 2026.
Visible fallback has been the stated intent since at least June, when Anthropic’s developer account extended it to Fable 5’s frontier-LLM safeguards.
The silent cases are the ones that should change your logging. Both are single user reports, and neither had a vendor response as of Sep 28.
anthropics/claude-code #97687, opened Sep 27 and still open with no comments, is the report in the opening. Its title: Sub-agents requesting model: "opus" silently continue on claude-opus-4-8 after a cyber-content refusal — parent report still says “Opus”. The reporter ran in-process subagents on Claude Code 2.1.282 and 2.1.283 with switchModelsOnFlag left at its default, and writes that the fallback itself matches the documented behavior. The undocumented part is propagation to the parent: “I only found the switch by reading message.model line-by-line in the sub-agent’s own transcript JSONL.”
Screenshot: GitHub, “Sub-agents requesting model: "opus" silently continue on claude-opus-4-8 after a cyber-content refusal…” (Sep 27, 2026), captured Sep 28, 2026.
Two days before that report, a practitioner posted the same shape in public: a disappointing night, then the discovery of where the subagents had actually run.
The second silent case is on OpenAI’s side. openai/codex #44637, opened Sep 10, is titled “Silent model reroute after a safety check: requests served by gpt-5.5 while the app kept showing gpt-6-astra”. It concerns the Codex app (26.903.71938) on ChatGPT Pro, and its only comment is a bot flagging a possible duplicate. One line in it should worry any fleet operator: “Nothing in the app or on disk records which model served a turn, so a user has no way to check this afterwards.”
A safety fallback is the fourth model swap, decided mid-session by a classifier
Operators already track three swaps. A router you chose picks a model per request, and which routed model shipped a change is its own audit question. A vendor retires a model and forwards its traffic, which forced-routing continuity covers. An upgrade moves a default, which is why the CLI upgrade canary checks the served default after every upgrade.
A safety fallback is the fourth swap, decided mid-session by a safety classifier reading your prompt and whatever the agent pulled in.
In a chat window the label is enough, because a person reads every answer. Chatbots suggest; agents act, and in a fleet nobody reads every turn. The code lands in a commit, the commit lands in a pull request, and the reviewer sees a diff with no model on it.
Claude Code’s model configuration page puts the next part plainly: “After a fallback, the session continues on the fallback model”. A run that trips once in its first hour can spend the rest of the night there.
Security lanes are the likeliest to trip the cyber classifier, and the help center adds that “Opus 5.5 isn’t currently available in the Cyber Verification Program.” Billing follows the answer too: a fallback response is billed at the responding model’s rates, with a cache-miss credit.
Even the vendor’s benchmark numbers carry fallback turns. The launch page’s footnote says safeguards stayed on during evals, with interventions completed by Opus 4.8 for cyber and by Opus 5 for biology and frontier-LLM work. If you promote a model to a default on evidence, check whether your own evidence did the same.
Step 1: Write down which model may answer for which
You can’t flag a mismatch without an expected set. For each model a lane requests, list the models that may legitimately answer in its place and the categories that route there. Take it from the help center rather than memory, and date it; the map below is the Sep 24 version. The date matters: on Sep 28, Claude Code’s model-config page added Sonnet 5.5, whose cyber flags re-run on Sonnet 5 and whose biology flags end in a refusal.
The documented map as of the help center’s Sep 24 update. Claude Code’s model-config lists the same cyber and biology targets for Fable 5.1 and Fable 5.
Three details from the Claude Code docs belong in the same register. Category-based fallback needs Claude Code v2.1.219 or later (Jul 24, 2026), so record the client version beside the map. On Opus 5, a biology-flagged request ends in a refusal, not a fallback. And on Amazon Bedrock, Google Cloud’s Agent Platform and Microsoft Foundry, fallback targets resolve through your deployment or ANTHROPIC_DEFAULT_OPUS_MODEL; if either model can’t be identified there, no switch happens and the refusal stands.
The map does two more jobs later. It is the triage key for your review queue: a cyber fallback on a lane that writes detection rules is expected, while a biology fallback on a billing service is a question somebody should ask. It is also the menu for the fixture in step 6.
Step 2: Log requested and served model on every call
One row per model call, whatever the surface, written by your wrapper rather than by the agent. These are the fields that carry the weight:
| Field | Source | Why it is there |
|---|---|---|
ts, lane, session_id, call_id |
your wrapper | Joins the row to the run, the commit and the weekly count |
requested_model |
request body or lane config, alias resolved at run start | Intent. Resolve aliases once, or every opus call reads as a mismatch |
served_model |
API: top-level model. Subagent: its own transcript |
The fact the whole log exists to capture |
iterations |
usage.iterations: a message entry per declining hop, a fallback_message entry for the server |
Records each hop inside one call |
fallback |
the fallback content block, from and to, when present |
Names both models on the turn that declined; absent on sticky turns, so never the key |
stop_reason, category |
stop_reason: "refusal", stop_details.category |
Separates a refusal, where nobody answered, from a fallback, where someone else did |
resolved_model, models_used |
Claude Code PostToolUse on a foreground Agent call |
The harness’s own record for foreground subagents |
agent_transcript, commits, client_version |
SubagentStop hook input, git, the CLI |
Where to re-read a background subagent, what to tag, and which fallback rules applied |
Key the check on served_model and iterations, never on the fallback block. The API refusals and fallback docs describe sticky routing: for about an hour, org-scoped and best-effort, later turns of a conversation that fell back go straight to the fallback model. “Such a turn carries no fallback content block, because no model declined that turn.”
A log that waits for the block sees the first swapped turn and misses every turn after it. The top-level field has no such gap: “The top-level model field reports the model that produced the returned message, whether that is the requested model or a fallback.”
Two API details go in the lane config. Server-side fallback is beta (header server-side-fallback-2026-07-01), works on the Claude API only, and isn’t available on Message Batches, Bedrock, Google Cloud or Foundry, where the documented routes are SDK middleware or a manual retry. Unconfigured, a flagged request comes back as HTTP 200 with stop_reason: "refusal" and a category of cyber, bio, frontier_llm, reasoning_extraction, general_harms or null. Whichever route a lane takes, the top-level model answers the question.
Step 3: Read the served model where it was recorded, never from the parent’s summary
A parent agent’s summary is a claim about its children, written by a model that may not know they were swapped. Read the served model at the layer that saw the response. For API lanes, that is the response.
For Claude Code subagents, the hooks reference documents two places. A PostToolUse hook on a foreground Agent call receives tool_response.resolvedModel, the model the subagent started on, “which may differ from the requested model”, and modelsUsed: “Models used in order, with consecutive repeats collapsed; set only when the model was swapped mid-run. Requires Claude Code v2.1.212 or later”. The docs don’t say whether a safety fallback populates modelsUsed, so test it before you rely on it.
Background subagents are the harder case, and they have been the default since v2.1.198 (Jul 1, 2026). A backgrounded call returns status: "async_launched", and resolvedModel names the model in use when the agent moved to the background; a later swap isn’t in the hook. The durable record is the subagent’s own transcript, whose path a SubagentStop hook receives as agent_transcript_path. The per-message model sitting in message.model comes from the #97687 report, not a documented schema, so parse it defensively and alert when the field is missing.
| Surface | What the docs or reports say | Labeled or silent |
|---|---|---|
| Claude apps (chat) | Notice, plus a label naming the model that answered; picker stays on the fallback | Labeled |
| Claude Code, interactive | Notice in the transcript; the session continues on the fallback model; a PostModelSwitch hook (v2.1.251+) fires with source: "auto" and can’t block |
Labeled, in the transcript |
| Claude Code subagent, in-process | Fallback matched the docs; nothing reached the parent (user report, #97687) | Silent to the parent |
Claude Code headless, -p |
With switchModelsOnFlag: false, the flagged request ends as an error |
Only as loud as the transcript you read |
| Claude API, fallback configured (beta) | fallback block on the declining turn; top-level model on every turn; sticky turns carry no block |
Labeled in the payload, if you parse it |
| Claude API, default | HTTP 200 with stop_reason: "refusal"; no swap |
A refusal, not a swap |
| Codex app | Selector kept showing gpt-6-astra while the usage page showed gpt-5.5 (user report, #44637) | Silent, per the report |
| OpenAI API | Error code cyber_policy; access may be temporarily revoked |
An error, not a swap |
For a surface that records nothing, the served-model column reads unknown, and a lane that needs provenance treats unknown as a mismatch. Reconcile those lanes weekly against the vendor’s usage page, which is where the #44637 reporter found the swap.
Step 4: Compare, alert, tag and re-review in one pass
One comparison, three outcomes. The parent’s summary is not an input.
Run the comparison as rows land, not in a nightly batch. A mismatch at 01:10 should page before the run writes forty more commits on the fallback model. The check below is illustrative and assumes your wrapper logs one JSON object per call, holding the requested model and the raw response.
# served_model_check.py (illustrative): flag calls where another model answered
import json, sys
RESOLVED = {"opus": "claude-opus-5-5"} # alias -> model ID it resolved to at run start
for line in open(sys.argv[1], encoding="utf-8"):
call = json.loads(line)
resp = call.get("response") or {}
want = RESOLVED.get(call["requested_model"], call["requested_model"])
got = resp.get("model") # top-level field: present on sticky turns too
hops = [i.get("type") for i in (resp.get("usage") or {}).get("iterations", [])]
if got != want or "fallback_message" in hops:
print(json.dumps({
"ts": call["ts"], "lane": call["lane"], "call_id": call["call_id"],
"requested": want, "served": got or "unknown", "hops": hops,
"category": (resp.get("stop_details") or {}).get("category"),
"action": "alert, tag commits, re-review",
}))
For a subagent, the same question takes one illustrative line over its transcript, using the field name from the #97687 report:
jq -r 'select(.message) | .message.model // "MISSING"' "$AGENT_TRANSCRIPT" | uniq -c
A flagged row does three things. It alerts the lane owner in the channel where failed builds go. It tags the commits from that turn onward, with a trailer such as Served-Model: claude-opus-4-8 and a served-model-mismatch label on the pull request; TASKS.md receipts make this automatic by putting the served model in every tick. And it queues a re-review before merge: re-run the tests, then have a reviewer read the diff knowing which model wrote it.
That isn’t blame. The help center calls the fallback “the less capable model”, and your approval covered the other one.
Step 5: Fail closed where change control names the model
Some lanes can live with a logged fallback. Others can’t: security tooling whose review assumed a specific model, regulated code, anything whose change record names the model behind each change. For those, Claude Code documents two levers, and both belong in settings, not in a prompt.
The first is switchModelsOnFlag, which defaults to true and can live in any settings file. Per the settings reference, false makes an interactive session pause and ask, and “where no dialog can show, such as a -p run, the flagged request ends as an error”.
The second is the allowlist: “The fallback model is checked against availableModels. When it is blocked, no fallback occurs.” The refusal stands, and your wrapper records a failed task instead of a quiet success. Put the list in managed settings, where it alone applies; user, project and local lists concatenate, so an entry in any of them can re-admit the fallback target. Entries can name a family such as opus, a version prefix such as claude-opus-5-5 or a full model ID, and a bare opus would admit Opus 4.8 too, so the illustrative shape below uses the prefix:
{
"switchModelsOnFlag": false,
"availableModels": ["claude-opus-5-5"]
}
Know where these gates stop. They govern Claude Code, not raw API calls; there, a strict lane stays strict by never configuring fallbacks and treating stop_reason: "refusal" as a failed task. The docs don’t say whether the setting reaches in-process subagents, and the #97687 reporter asks Anthropic to document exactly that, so prove it with the fixture before trusting it. When a gate fails anyway, the log is the wall behind it: any mismatched row on a strict lane blocks merge until a person clears it or the task re-runs on the requested model.
Step 6: Count fallback turns weekly, and keep one fixture that proves the log
Every Monday, roll the log up per lane: fallback turns over total turns, by category and by served model. A lane that jumps from zero changed its inputs, or the classifiers changed; find out which. A lane living on the fallback model has made a model decision nobody signed, so approve it or fail it closed. The count feeds cost as well: pre-output refusals in bio, frontier_llm and reasoning_extraction are billed “as of September 2026”, and cyber and general_harms ones aren’t.
Then prove the plumbing with two fixtures. A recorded fixture runs in CI on every change to the wrapper: three saved responses, one whose model differs from the request, one with a fallback block and one sticky turn without it. The check must flag all three.
A live fixture runs monthly and after each CLI upgrade: one request from your own log that already triggered a documented fallback on an Anthropic model, replayed once in a sandbox with no write access. Don’t craft new triggers, and keep the volume at one. No source says whether deliberate test triggers count against the “account trust signals” the help center says Anthropic weighs.
Never point either fixture at OpenAI’s API. Its cyber checks don’t reroute: “API requests will return an error with the error code cyber_policy”, and without a per-user safety_identifier, “access may be temporarily revoked for the entire organization”. A live fixture that stops tripping is a finding too. Either the classifiers moved or your log did.
Safety fallbacks that slip past the log, and the signal for each
Most misses come from reading the wrong field or the wrong layer. Each has a signal you can alert on.
| Failure | Signal | Fix |
|---|---|---|
| The sticky turn | served_model differs on rows with no fallback block |
Key on the top-level model field (step 2) |
| A background subagent swaps after launch | The hook’s resolvedModel matches; the transcript shows a second model |
Read agent_transcript_path at SubagentStop for every background run |
| The confident parent | A summary credits one model while a child transcript shows two | Never feed summaries into the log; this is #97687 in one row |
| The instruction file trips the classifier | A lane falls back on turn one, every session | Run the lane with claude --safe-mode, which the model-config page names for checking whether CLAUDE.md content triggers it |
| The alias moved | Every opus row flags at once, across lanes |
Not a fallback: Claude Code 2.1.280 changed what the alias resolves to. Re-resolve at run start |
| The fallback model is busy | Refusals carry stop_details.recommended_model |
Route by lane policy, never automatically to the recommendation |
Provenance is a fleet record, not a model label
A label in one vendor’s window answers who answered in that window. Your fleet runs several CLIs, the raw API and subagents that report through a parent, so the answer has to live in the layer that runs all of them: the wrapper that logs each call, the policy that fails a strict lane closed, the Monday count. That layer is what a multi-agent command center is for, and the same log is what fleet replay needs to show which model took each step.
FAQ
Why did Claude switch to Opus 4.8 in the middle of my session?
A safety classifier flagged a request as potentially higher-risk offensive cybersecurity work, so Opus 5 or Opus 5.5 handed the turn to Opus 4.8. The picker then stays on the fallback for the rest of the conversation. Switch back from the model picker, or turn off automatic switching to be asked first.
Does the Claude API fall back to another model automatically?
No. On the API, automatic switching is off until you configure it, and a flagged request returns HTTP 200 with a refusal stop reason. Server-side fallback is an opt-in beta on the Claude API only, not Message Batches, Bedrock, Google Cloud or Foundry. Log the top-level model field either way.
Sources
- Claude Support: Why Claude switched models in your conversation with Opus 5 or Opus 5.5 (updated Sep 24, 2026)
- Anthropic: Claude Opus 5.5 launch page (Sep 22, 2026)
- Claude API: Refusals and fallback
- Claude Code: Model configuration, “Automatic model fallback”
- Claude Code: Settings reference,
switchModelsOnFlag - Claude Code: Hooks reference, Agent
tool_responseandSubagentStop - Claude Code changelog (2.1.198, 2.1.212, 2.1.219)
- anthropics/claude-code #97687 (user report, Sep 27, 2026)
- openai/codex #44637 (user report, Sep 10, 2026)
- OpenAI API: Cybersecurity safety checks
