The Fast-Worker Lane Contract: Swap the Fast AI Coding Model, Keep the Lane
Fast AI coding models turn over in weeks. Write a lane contract for your cheap subagent worker, drill the swap on fixtures, and change one config line.
Go deeper. Build your own.
On September 10, GitHub retired MAI-Code-1-Flash across Copilot Chat, inline edits, ask and agent modes, and code completions, with one line for everyone who had built on it: “Please update your workflows and integrations to use a supported model.” For a team whose cheap worker sat in a lane with a written contract, that notice meant one config change and an afternoon of fixtures. For a team whose prompts, review habits and stop rules had grown around one model, it meant a rewrite.
Expect the notice again. The fast AI coding model under your subagents is the most replaceable part of your stack: in the three weeks after GitHub’s notice, OpenAI, Xiaomi, MiniMax, Anthropic and inclusionAI all shipped or listed a new small or fast model. Their list prices differ by more than an order of magnitude, and “Flash” in a model name tells you nothing about which end of that range you’re on.
So write the contract for the lane, not the model: what the worker owns, what it never touches, when it hands back to the main model, and what stays fixed when you swap it. Then rehearse the swap before a vendor schedules it for you.
Five fast models in two weeks, and one forced retirement
GitHub’s deprecation notice is short and final: MAI-Code-1-Flash came out of every Copilot surface at once. Its replacement had been sitting there since August 11, when GitHub’s MAI-Code-1.1-Flash changelog listed a 73% lower list price and added vision.
Screenshot: GitHub Changelog, “MAI-Code-1-Flash deprecated” (Sep 10, 2026), captured Oct 5, 2026.
Then the fast tier turned over. Xiaomi’s MiMo-V2.6-Flash reached OpenRouter on September 21 at $0.10 input and $0.28 output per million tokens. OpenAI launched GPT-6 Sol and Luna on September 22, with Luna at $0.10 and $0.50 against GPT-5.6 Luna’s $0.20 and $1.20, and said Luna “improves on its predecessor by 5.4 percentage points at 58% lower cost per task” on its AutomationBench.
MiniMax put M3.1-Flash-Preview inside its own coding tool on September 27 (DataNorth), and its model docs say it “is available only through M Plan and MiniMax Code for now”, so there is no pay-as-you-go API price yet. Anthropic shipped Claude Sonnet 5.5 on September 28 at $2 and $10, “30%+ faster than Sonnet 5”. inclusionAI’s Ling 3.1 Flash appeared on OpenRouter at $0 on October 2.
Screenshot: OpenAI, “Introducing GPT-6 Sol and Luna” (Sep 22, 2026), captured Oct 5, 2026.
Claude Haiku 5.5 is not on that list: Sonnet 5.5’s launch page says it “will join the Claude 5.5 family in the coming weeks”, and on October 5 the Claude models overview still lists Claude Haiku 4.5 as the current Haiku. Haiku 5.5 is announced, not shipped, so a plan that assumed a new Haiku is waiting on a release.
Usage moved as fast as the launches. OpenRouter’s rankings for the week of September 28 show GPT-6 Luna up 116% in weekly tokens while GPT-5.6 Luna fell 49%, with three separate DeepSeek Flash builds in the top ten at once. Workers get replaced in weeks, so the lane has to outlive any one of them.
One more fact surprises people who think they already run a cheap worker. Claude Code’s built-in Explore subagent runs on “the main conversation’s model”, and the subagents docs tell you to “define one with model: haiku” if you want exploration on a lower-cost model. On a default Claude Code install on a paid plan today, your exploration runs on Opus.
Fast workers no longer only autocomplete. In agent mode they read the repo, run tests and write files, so a model swap changes the behavior of something with write access, and that calls for a contract and a drill.
The lane-contract runbook: nine steps to a worker you can swap in an afternoon
1. Name what the lane is, and what it isn’t
A fast-worker lane is a standing assignment: a class of work that always goes to a cheaper, faster model under a contract that never names the model. Three neighbours get confused with it.
- It is not
/fast. Fast mode runs the same model faster at a higher price and belongs in its own speed budget. A fast-worker lane runs a different, cheaper model. - It is not a router. Routing as code picks a tier per request with written criteria and a confidence floor, as in routing cheap models with Jev. Borrow that piece’s escalation floor; this lane is the standing assignment underneath it.
- It is not a model runbook. One model’s pricing quirks, such as off-peak windows and cache-hit rates, belong in their own page, the way DeepSeek V4.1-Flash’s off-peak runbook does.
Write the lane’s name, owner and main model at the top of the contract. The worker model goes in one place only: the config line in step 6.
2. Write what the worker owns
List task classes, and give each one a “done” line. A fast worker that doesn’t know when it has finished spends its savings on extra turns.
| Task class | Done means | Example |
|---|---|---|
| Read-only exploration | file paths plus a summary, no edits | “where is retry logic configured” |
| Grep and summarize | findings with file and line references | every call site of a deprecated helper |
| Test-log triage | failing test named, first error quoted, suspect file | a red CI run on a feature branch |
| Lint and format fixes | linter exits clean, no logic in the diff | an import-order sweep |
| Single-file mechanical edits | one file changed, existing tests pass | renaming a parameter inside one module |
| Doc drafts | a draft on a branch, reviewed by a human | a README section for a new flag |
Illustrative ownership table: the classes come from the contract, the examples are yours to replace.
Keep the list short on purpose. A class that is not on it goes to the main model, which is the cheapest mistake the lane can make. Review the list whenever the lane changes owner; a new owner inherits the contract, not the habits.
3. Write what it never owns
These go straight to the main model, with no attempt by the worker first:
- Authentication, authorization and anything touching secrets.
- Payments and billing code.
- Database migrations.
- Multi-file refactors.
- Any change to code without a test that would catch the worker being wrong.
- Dependency upgrades and CI configuration, because their blast radius is the whole repo.
- Any model you cannot attribute to a named vendor with published terms, however well it ranks this week.
If a class is ambiguous, it belongs on the never-owns list until a fixture proves otherwise.
4. Write the escalation triggers
Escalation is the contract’s safety valve, and it has to fire on facts, not on the worker’s confidence:
- Two failed attempts on the same task.
- A diff that touches more files than the class allows: one for mechanical edits, whatever you set for lint sweeps.
- A test fails that the worker didn’t write.
- The worker asks a question.
On escalation the main model receives the worker’s transcript and diff, not a fresh prompt, so the expensive model starts where the cheap one stopped. Count escalations per day per lane; that number drives the kill switch in step 9.
Start each file limit at the class’s own size and raise it only after a drill shows the worker stays inside the review gate at the larger size. Limits that drift upward one exception at a time are how a mechanical-edit lane ends up doing refactors nobody assigned it.
5. Freeze what stays fixed across swaps
Five things never change in the same commit as the worker model: the prompt files, the tool allowlist, the review gate, the stop rule (maximum turns and tokens per task) and the eval fixtures. All five live in the repo under version control. The model ID is the only variable.
The lane contract on one page: four boxes that never mention a model, and a swap drill that changes exactly one line.
This is where most teams slip. A new worker behaves a little differently, someone tunes a prompt to suit it, and now the next swap has two variables. Tune prompts in their own change, against the incumbent, with fixtures.
Keep the whole contract in one file next to the prompts, so a reviewer sees it in the same diff as any change to the lane. An illustrative shape:
lane: fast-worker
owner: platform-team
main_model: opus
worker_model: haiku
owns: [explore, grep-summarize, test-log-triage, lint-format, single-file-edit, doc-draft]
never_owns: [auth, secrets, payments, migrations, multi-file-refactor, untested-code, deps, ci-config]
escalate_when: {failed_attempts: 2, max_files: 1, foreign_test_failure: true, worker_question: true}
fixed: {prompts: prompts/fast-worker/, allowlist: [Read, Grep, Glob, Edit], review_gate: required, max_turns: 12}
kill_switch: {escalation_rate_over: 2x_baseline, window: 24h, fallback: main_model}
6. Wire the worker as one config line
Every runtime has a single place to name the worker. Use it and nothing else.
In Claude Code, CLAUDE_CODE_SUBAGENT_MODEL sets the model for subagents, and a subagent definition’s model: field sets it per agent. Because the built-in Explore subagent follows the main model, define your own exploration agent for the lane. An illustrative definition, with the worker in the model: line:
---
name: explore-worker
description: Read-only exploration for the fast-worker lane. Returns file paths and a short summary.
model: haiku
tools: Read, Grep, Glob
---
Follow docs/lanes/fast-worker.md. Two failed attempts or any question: stop and hand back.
In Codex the equivalent is the model key in its config file, for example model = "gpt-6-luna" (illustrative).
Whatever the runtime, log the served model ID with every task result. A vendor can change what sits behind a name without a release note you’ll see, and a fixture baseline is only comparable if you know which model produced it. How that worker setting sits next to the other per-session knobs on a shared login is covered in one login, many lanes.
7. Build the fixture set from the lane’s own history
Fixtures are what make a swap boring. Build them once, from work the lane actually did:
- Pull 20 to 30 real tasks from the last month, at least three per owned class, each with the expected result written down.
- Add five tasks that should escalate: two from the never-owns list and three that trip each trigger in step 4. A worker that completes those is a failure, not a win.
- Record per run: completed (passed the review gate), escalated, attempts, tokens, cost and wall time.
- Score the lane on cost per completed task and escalation rate. The denominator matters; cost per completed task explains why the token price alone misleads.
Store the fixture results with the model ID and date, so the next drill has a baseline to beat.
Keep the set representative and dull. Tasks the lane never sees in practice make a candidate look better or worse for reasons that won’t hold in production. Refresh about a third of the set each quarter from new lane history, and never tune a prompt against the fixtures you use to judge a swap, or the drill grades its own homework.
8. Run the swap drill
Run it when a vendor retires or renames your worker (MAI-Code-1-Flash to 1.1-Flash, GPT-5.6 Luna to GPT-6 Luna), when a cheaper candidate appears, and once a quarter regardless.
Start with a shortlist, and use price only as a filter. The names don’t sort the models: MiMo-V2.6-Flash and Gemini 3.8 Flash share a suffix, and Gemini’s introductory input price is more than seven times MiMo’s or Luna’s.
List prices, not costs per task: a shortlist filter for the swap drill. Gemini 3.8 Flash shows its introductory input price only.
Google’s Gemini API pricing page lists Gemini 3.8 Flash at $0.75 per million input tokens through December 31, 2026, then $1.50, with thinking tokens billed as output. DataNorth’s launch write-up adds a warning that it “may use more reasoning tokens than 3.7 Flash”. That warning is the reason the fixtures decide, not the price table.
Treat speed the same way. The only per-model numbers in the window are OpenRouter’s own measurements: about 108 tokens per second and a 1.72-second median latency from Luna’s fastest provider, about 85 tokens per second from MiMo’s. Those vary by provider and by day, so measure wall time in your fixtures.
Then the drill itself:
- Run the fixture set on the incumbent and the candidate, with identical prompts, allowlist, review gate and stop rule.
- Compare cost per completed task and escalation rate. The candidate wins only if it is cheaper per completed task and its escalation rate is no worse.
- Change the one config line, in its own commit, with the fixture results in the commit message.
- Canary the candidate on a share of the lane’s tasks for a week, then promote it with the evidence a default deserves, as in the evidence that earns a model the default slot.
If a vendor retires your worker before the drill finishes, point the config line at the main model for the gap. Paying more for a week beats running an unmeasured worker.
9. Set the kill switch and the deprecation watch
The kill switch is one rule: if the lane’s escalation rate stays above your threshold for a full day, the config line reverts to the main model. Set the threshold from the fixture baseline, for example twice the escalation rate the worker showed in its last drill. Automate the revert if your runner can, and name the person who does it by hand if not.
Test the switch the way you test a backup. Once a quarter, flip the lane to the main model on purpose, confirm work still flows and the escalation counter resets, then flip it back. A kill switch nobody has pulled is a guess.
The deprecation watch is a subscription. Follow the changelog of every vendor whose model fills a worker slot, and treat a deprecation notice as the start of the swap drill that same day, not the day the model stops answering.
Where a fast-worker lane breaks, and the signal for each
- Prompts drift toward the new worker. Signal: a pull request that changes the model line and a prompt file together. Fix: split it, and rerun fixtures on the prompt change against the incumbent.
- Exploration never left the main model. Signal: subagent usage attributed to Explore at main-model rates, or no custom exploration agent in the repo. Fix: step 6.
- A “Flash” swap raises the bill. Signal: cost per completed task goes up after a swap to a model with a fast-sounding name. Fix: step 8’s price filter and fixtures.
- Escalation creep. Signal: the escalation rate climbs week over week with no change in task mix. Fix: check whether the vendor changed the model behind the same name; if the fixtures confirm it, run the drill.
- The worker over-reaches. Signal: merged worker diffs that touch more files than its class allows, or that edit code with no test. Fix: tighten the review gate to reject those diffs mechanically.
- The notice lands and nobody has a baseline. Signal: a deprecation notice and no stored fixture results. Fix: route the lane to the main model now, then build fixtures before choosing anything.
A lane contract is how a fleet survives model churn
An operating layer exists so that changes in one vendor don’t ripple through every workflow you run. The lane contract is that idea at the smallest useful size: the worker is a rented part, and the contract, the fixtures and the review gate are yours. That is the split AgentOps argues for at fleet scale, and the same one the loop-ownership register applies to prompts, tools and transcripts.
The three weeks after September 10 retired one fast worker and brought five new ones; a team with lane contracts needed one drill per lane and one changed line each. I’d rather spend that afternoon than the week a rewrite takes.
FAQ
What is the best cheap AI coding model for subagents right now?
There isn’t a stable answer. Between September 10 and October 2 the fast tier gained GPT-6 Luna, MiMo-V2.6-Flash, Sonnet 5.5 and more, while GitHub retired MAI-Code-1-Flash. Shortlist by list price, then run your own fixtures and pick by cost per completed task and escalation rate.
Is Claude Code /fast the same as using a fast AI coding model?
No. /fast runs the same model faster at a higher price, so it belongs in a speed budget. A fast-worker lane hands a defined class of work to a different, cheaper model under a written contract. One buys latency on the main model; the other moves scoped work off it entirely.
Does Claude Code’s Explore subagent use Haiku?
Not by default. Claude Code’s docs say the built-in Explore subagent runs on the main conversation’s model, which is Opus on most paid plans today. To run exploration on a cheaper model, define your own exploration subagent with model: haiku and point your lane’s exploration tasks at it.
Sources
- GitHub Changelog, “MAI-Code-1-Flash deprecated” (Sep 10, 2026)
- GitHub Changelog, “MAI-Code-1.1-Flash available in GitHub Copilot” (Aug 11, 2026)
- OpenAI, “Introducing GPT-6 Sol and Luna” (Sep 22, 2026)
- OpenRouter, GPT-6 Luna model page (released Sep 22, 2026)
- OpenRouter, MiMo-V2.6-Flash model page (Sep 21, 2026)
- Anthropic, “Claude Sonnet 5.5” (Sep 28, 2026)
- Claude Platform docs, “Models overview” (fetched Oct 5, 2026)
- Claude Code docs, “Subagents” (fetched Oct 5, 2026)
- OpenRouter, LLM Rankings (week of Sep 28, 2026)
- Google, Gemini API pricing (fetched Oct 5, 2026)
- MiniMax API docs, “Models” (fetched Oct 5, 2026)
- DataNorth, “Google launches Gemini 3.8 Flash & Gemini 3.8 Flash Cyber” (Sep 2, 2026; secondary)
