GPT-6.1 Sol vs Claude Sonnet 5.5: Same List Price, Different Agent Bills. Price the Workload Shape
Both list at $2/$10 with $0.10 cache reads, yet a 50-turn agent loop past 272K costs 29% more on one. Price your workload shape with this five-row worksheet.
Go deeper. Build your own.
Run a 50-turn agent loop at a steady 100K tokens of context and the two price sheets agree to the cent: $2.10 on GPT-6.1 Sol, $2.10 on Claude Sonnet 5.5. Let the same loop grow by 5K tokens a turn until it ends at 345K and the bills split, $3.51 against $2.71 (illustrative, at list prices). Nothing about the per-token sticker changed between those two runs. The shape of the work did.
That is the whole GPT-6.1 Sol vs Claude Sonnet 5.5 pricing story for anyone running agents rather than chats. Both models list at $2 input and $10 output per million tokens, both charge $0.10 for a cache read, and one of them reprices an entire request once its input crosses 272K tokens. A worksheet that prices your actual loop shapes on both sheets tells you which lanes are exposed. A sticker comparison tells you nothing.
This piece gives you that worksheet with five shapes filled in, the arithmetic behind each row, an alarm to set before the long-context line, and the recompute rule for the day either page changes. It does not pick a model. Model quality, tokenizer efficiency and task success sit outside a price sheet, and cost per completed task is still the number that settles a lane.
The two price sheets as read on Oct 7: $2/$10, $0.10 reads and one 272K line
Anthropic released Claude Sonnet 5.5 on September 28, 2026, per its announcement, with launch coverage such as ITdaily’s the next day. OpenAI followed on September 29 at DevDay with GPT-6.1 Sol, per DataCamp’s launch write-up; OpenAI’s launch post reportedly carries the same date but refused our fetch. The API model is gpt-6.1-sol. It is not the “GPT-6 Sol” that ChatGPT now uses as its chat default, and prices for one say nothing about the other.
We read both pricing pages on the evening of October 7, US time (about 00:20 to 00:45 UTC on October 8), and re-read them before filing. OpenAI’s pricing page lists gpt-6.1-sol at $2.00 input, $0.10 cached input, $2.50 cache writes and $10.00 output per million tokens for short context. Its long-context columns for the same row read $4.00, $0.20, $5.00 and $15.00.
The page defines the line exactly: “Short context: ≤272K input tokens. Long context: >272K input tokens.” The GPT-6.1 Sol model page adds that the long-context rate applies to the full request, and lists a 1,050,000-token context window with 922,000 maximum input tokens and 128,000 output.
Screenshot: OpenAI Developers, “Pricing | OpenAI API” (flagship models table, Oct 7, 2026), captured Oct 7, 2026.
Anthropic’s pricing documentation lists Claude Sonnet 5.5 at $2 input, $2.50 for a 5-minute cache write, $4 for a 1-hour cache write, $0.10 for cache hits and refreshes, and $10 output per million tokens. Its footnote prices Sonnet 5.5 cache hits at 0.05x base input. The consumer-facing claude.com pricing page shows the same $2, $10, $2.50 and $0.10.
Screenshot: Claude Platform Docs, “Pricing - Claude Platform Docs” (model pricing table, Oct 7, 2026), captured Oct 7, 2026.
The long-context section is where the sheets part ways. Anthropic says Claude 4.6 and later models, except Haiku 5.5, “include the full 1M token context window at standard pricing,” and spells it out: “A 900k-token request is billed at the same per-token rate as a 9k-token request.” Batch is 50% off on both vendors, and Anthropic’s caching multipliers stack with its batch discount.
Screenshot: Claude Platform Docs, “Pricing - Claude Platform Docs” (long context pricing passage, Oct 7, 2026), captured Oct 7, 2026.
Second, OpenAI’s table shows a $2.50 cache-writes column for gpt-6.1-sol, and its prompt-caching guide says caching is on by default and that cache writes on GPT-5.6 and later cost 1.25× the uncached input rate. Our worksheet charges every new token as a write. Tokens placed after the last explicit cache breakpoint bill at plain input instead; if every new token billed that way, the flat-100K loop would cost $1.975 instead of $2.10. Third, the same guide keeps a cached prefix for 30 minutes after its last write or reuse, against Anthropic’s five-minute default, so the pause scenario below is priced for Anthropic only.
Why an agent loop exposes what one chat request hides
A single chat turn at 20K tokens bills about the same anywhere with a $2/$10 sticker. An agent loop is different in two ways that the sticker hides. It re-sends its whole history every turn, so most of what you pay for is cache reads, and its history grows, so a lane that starts at 100K can finish well past 272K without anyone deciding it should. Both effects compound per turn, and both are invisible until the invoice arrives.
Agents also pause. A lane waiting five minutes for a human approval can come back to an expired cache and pay to write its entire prefix again. A one-vendor cache meter catches that inside a single sheet; the worksheet below catches it across two.
Build the workload-shape worksheet in six steps
Step 1: Log four numbers per turn for every lane
You cannot price a shape you have not measured. For each agent lane, record per turn the input tokens sent, how many of those were cache reads, how many were cache writes, and the output tokens, reasoning included. Add the wall-clock gap since the previous turn, because that gap decides whether the cache survived. One line per turn is enough; the field names below are illustrative.
{"lane": "repo-refactor", "vendor": "openai", "model": "gpt-6.1-sol",
"turn": 36, "input_tokens": 275000, "cache_read": 270000, "cache_write": 5000,
"output_tokens": 2000, "gap_since_prev_s": 41, "tier_billed": "long",
"sheet_read": "2026-10-07"}
Measure on both vendors when a lane runs on both. Anthropic’s own pricing page notes that its current tokenizer “produces approximately 30% more tokens for the same text” than its previous one, so the same transcript is not the same token count across tokenizers. The worksheet below assumes equal counts to isolate the price mechanics. Your log should replace that assumption with measured counts.
Step 2: Sort lanes into five workload shapes
Most agent lanes fall into one of five shapes. Name the shape in the lane’s config so the bill has a reason attached.
Classify from a week of logs, not from the lane’s intent. A lane described as “quick triage” that keeps every tool result in history is a growing lane, whatever its owner calls it. Plot input tokens per turn for the ten longest runs; a flat line, a ramp and a step at load time are the three patterns that matter, and output share tells you whether the cache price even moves the total.
| Shape | What it looks like | Typical lane |
|---|---|---|
| Flat 100K × 50 turns | Context stays near 100K; old turns get compacted | Support triage with a fixed knowledge pack |
| Growing 100K → 345K × 50 turns | Each turn adds about 5K and nothing is dropped | Long refactor or research loop |
| Flat 300K × 50 turns | A large repo or document set loaded up front | Codebase-wide review |
| Mostly-cached 20K × 50 turns | Small context, output-heavy turns | Drafting, form filling, short tool calls |
| Batch overnight, 100 × 300K | Independent requests, no cache, results by morning | Document intake or bulk classification |
Step 3: Price every shape on both sheets
This is the artifact. Every number is illustrative: list prices as read on October 7 (US evening), standard tier unless the row says batch, no data-residency uplift, equal token counts on both vendors, and every new token per turn billed as a 5-minute cache write. The re-run column is the date you owe the next recompute.
| Workload shape | Input tokens (run) | Cached share | Output tokens | GPT-6.1 Sol bill | Sonnet 5.5 bill | Delta | Mechanic that caused it | Re-run date |
|---|---|---|---|---|---|---|---|---|
| Flat 100K × 50 turns | 5,000,000 | 95% | 100,000 | $2.10 | $2.10 | $0.00 | Both under 272K; identical rates | Next change to either pricing page |
| Growing 100K → 345K × 50 turns | 11,125,000 | 97.8% | 100,000 | $3.51 | $2.71 | +$0.80 (+29%) | Turns 36–50 pass 272K; OpenAI reprices the whole request | Next change to either pricing page |
| Flat 300K × 50 turns | 15,000,000 | 98.3% | 100,000 | $5.70 | $3.10 | +$2.60 (+84%) | Every turn past 272K: 2× reads and writes, 1.5× output | Next change to either pricing page |
| Mostly-cached 20K × 50 turns | 1,000,000 | 95% | 50,000 | $0.72 | $0.72 | $0.00 | Under the line; output is 69% of the bill | Next change to either pricing page |
| Batch overnight, 100 × 300K | 30,000,000 | 0% | 200,000 | $61.50 | $31.00 | +$30.50 (+98%) | Batch halves both sheets; the 272K line still applies on OpenAI | Next change to either pricing page |
The arithmetic for the rows that differ, so you can check it against your own reading of the pages:
Growing 100K -> 345K, 50 turns, 5K new + 2K out per turn (illustrative)
context at turn k = 100K + 5K(k-1); turn 35 = 270K, turn 36 = 275K
cache reads = 50 x 95K + 5K x 1,225 = 10,875,000; writes 250,000; output 100,000
Sonnet 5.5: 10,875,000 x 0.10 + 250,000 x 2.50 + 100,000 x 10 = 2.7125
GPT-6.1 Sol 1-35: 6,300,000 x 0.10 + 175,000 x 2.50 + 70,000 x 10 = 1.7675
GPT-6.1 Sol 36-50: 4,575,000 x 0.20 + 75,000 x 5.00 + 30,000 x 15 = 1.7400
GPT-6.1 Sol total = 3.5075 (+29%) (all prices USD per million tokens)
Flat 300K, 50 turns: 295K read + 5K write + 2K out per turn
Sonnet 5.5: 0.0295 + 0.0125 + 0.020 = 0.062 a turn -> 3.10
GPT-6.1 Sol: 0.0590 + 0.0250 + 0.030 = 0.114 a turn -> 5.70
Batch overnight, 100 requests x 300K input, 2K out, no cache
Sonnet 5.5 batch ($1 in / $5 out): 0.30 + 0.010 = 0.310 -> 31.00
GPT-6.1 Sol batch long ($2.00 in / $7.50 out): 0.60 + 0.015 = 0.615 -> 61.50
In the growing loop, the whole $0.80 gap sits in the last 15 turns. Sonnet 5.5 bills those turns $0.945 and GPT-6.1 Sol bills them $1.74. Turns 1 through 35 cost the same to the cent on both. The batch row uses the long-context batch rates OpenAI lists for this model ($2.00 input and $7.50 output once past 272K), which is why a 50% discount does not close the gap.
Read the two zero-delta rows as carefully as the others. The mostly-cached 20K lane spends 69 cents of every dollar on output, so the cache-read price barely moves it and the 272K line never comes near. Moving that lane between vendors on price grounds would be a waste of a migration. The flat-100K lane is equally safe until somebody removes its compaction step, at which point it becomes the growing shape without changing its name.
The batch row is the one most teams miss. Batch jobs feel cheap by construction, and nobody watches their per-request size because nobody watches them at all. Take a document-intake job whose inputs creep from 250K to 300K over a quarter: it crosses the line quietly, and its bill more than doubles on one vendor while its request count stays flat.
Illustrative: three 50-turn loop shapes priced from both list-price pages read Oct 7, 2026; equal token counts assumed.
Step 4: Put a context alarm at 80% of the vendor’s line
The cliff is a property of the vendor, not of your task, so the alarm belongs in the lane’s config next to the model ID. For a 272K line, 80% is about 218K. When a lane’s 95th-percentile input per turn crosses that, compact or summarize before the next turn, not after the bill shows it. A lane on a vendor with no long-context tier gets no alarm for this reason, though it may still want one for latency.
lane: repo-refactor
model: gpt-6.1-sol
price_sheet_read: 2026-10-07
long_context_line_tokens: 272000
alarm_at_tokens: 218000 # 80% of the line
on_alarm: compact_history # summarize oldest turns before the next request
p95_window_turns: 20
Every request passes the same check; the alarm at 218K gives the lane one chance to compact before the whole request reprices.
Compaction is not free either. A summary turn costs output tokens and can drop facts the agent later needs. Price it in the worksheet as its own row if a lane compacts often.
Step 5: Price the pauses against the cache lifetime
Anthropic’s 5-minute cache write lasts five minutes; the 1-hour write costs $4 per million tokens and lasts an hour. An agent that waits longer than five minutes for an approval comes back to a cold cache. In the flat-100K shape, that turn re-writes its whole 100K prefix at $2.50 per million, so it costs $0.25 plus $0.02 of output: $0.27 against the usual $0.042, about 6.4 times the turn price (illustrative).
A lane that pauses at every step should either buy the 1-hour write or keep its human gates short. Log approval-wait minutes beside each lane’s cache lifetime, and the worksheet will show you which one you are paying for.
OpenAI’s prompt-caching guide gives gpt-6.1-sol a 30-minute cache lifetime, renewed by each write or reuse, so the same pause costs nothing extra on that lane until it passes half an hour. Measure it anyway: send the same prefix after a known gap and check whether the response reports cached tokens.
Step 6: Recompute the worksheet when either page changes
Sonnet 5.5’s cache-read price halved nine days after launch. A worksheet filled in on September 30 had the wrong number in every Sonnet 5.5 row by October 7. Treat both pricing pages as inputs with a read date, not constants.
The recompute takes ten minutes if the worksheet keeps its inputs separate from its results. Store the eight gpt-6.1-sol cells and the five Sonnet 5.5 cells in one small table with the read date and URL, and let every row reference them. When a page changes, you update one table and every bill recalculates, including the deltas you showed someone last week.
The rule set, short enough to pin above the worksheet:
- Price the shape, not the sticker. Every lane has a named shape and a worksheet row; nobody compares vendors on the headline rate alone.
- Put a context alarm at 80% of each vendor’s long-context line, in the lane’s config.
- Log cache-hit share per lane per day; a drop means pauses, compaction or a prompt change is re-buying the cache.
- Recompute the worksheet whenever either pricing page changes, and stamp the read date on it.
- Keep promotional and introductory prices out of this worksheet; their expiry dates live in a separate price-expiry register.
If a lane bills in credits or a subscription allowance rather than API tokens, it does not belong here at all; translate it first with the meter dialects guide, and keep plan limits on the subscription plans side.
GPT-6.1 Sol vs Sonnet 5.5 bill failures, with the signal and the first fix
| What breaks | The signal you would see | First action |
|---|---|---|
| A growing lane drifts past 272K on GPT-6.1 Sol | Per-turn cost steps up mid-run while output per turn stays flat; usage shows long-context billing | Turn on the 218K alarm and compact; re-price the lane as the growing shape |
| A comparison deck still uses “$0.10 vs $0.20” | Sonnet 5.5 rows priced at twice the cache-read rate on today’s page | Re-read both pages, correct the rows, stamp the read date |
| Approval pauses re-buy the cache | Cache writes spike on the first turn after each human gate; cache-hit share falls | Log wait minutes per gate; test the 1-hour write on that lane |
| Equal token counts assumed across vendors | Measured tokens per task differ from the worksheet’s input column | Replace modeled counts with logged counts from Step 1 |
| A batch job sized under the line grows over it | Overnight batch bill roughly doubles with no change in request count | Check document sizes against 272K; split or chunk inputs before submission |
| OpenAI cache-write billing differs from the worksheet | Invoice lower or higher than the worksheet for flat lanes | Reconcile one day of usage against both readings; record which one the invoice matches |
Keep one worksheet for every model lane in the fleet
A fleet with lanes on two vendors has two price sheets, two context lines and two cache lifetimes moving underneath it. The worksheet belongs where the lanes are assigned, not in a spreadsheet one person maintains. When the mode-to-model assignment table moves a lane from one vendor to another, the worksheet row moves with it and gets re-priced on the new sheet the same day.
That is the job a multi-agent command center does for the fleet: one place that shows each lane’s model, its context trend against the vendor’s line, and the read date of the price sheet it was budgeted on. The sticker is the same on both vendors this week. The bills are whatever your shapes make them.
FAQ
Is GPT-6.1 Sol cheaper than Claude Sonnet 5.5?
On list price they tie: both are $2 input and $10 output per million tokens with $0.10 cache reads. They diverge on long context. Past 272K input tokens OpenAI bills the whole request at 2× input and cache rates and 1.5× output, while Sonnet 5.5 stays flat across its 1M window.
Do GPT-6.1 Sol and Claude Sonnet 5.5 charge the same for cache reads?
Yes, as of October 7, 2026, both list $0.10 per million cached tokens. Sonnet 5.5’s rate was halved from $0.20 that day in Anthropic’s Haiku 5.5 announcement. GPT-6.1 Sol’s cached rate rises to $0.20 once a request passes 272K input tokens. Re-read both pages before budgeting.
How do I estimate what an agent loop costs on each model?
Log input, cached, written and output tokens per turn, then price each turn on both sheets, applying OpenAI’s long-context rates to any request over 272K. Sum the turns. A 50-turn loop at 100K costs $2.10 on both; the same loop growing to 345K costs $3.51 versus $2.71 (illustrative).
Sources
- OpenAI, “Pricing | OpenAI API”, gpt-6.1-sol short and long context rows, batch rates (read Oct 7, 2026)
- OpenAI, GPT-6.1 Sol model page, context window and full-request long-context billing (read Oct 7, 2026)
- Anthropic, “Pricing - Claude Platform Docs”, Sonnet 5.5 row, cache multipliers, long context and batch (read Oct 7, 2026)
- Anthropic, Claude pricing page, Sonnet 5.5 API rates (read Oct 7, 2026)
- Anthropic, Claude Haiku 5.5 announcement, including the Sonnet 5.5 cache-read cut (Oct 7, 2026)
- DataCamp, GPT-6.1 Sol launch write-up, DevDay release date and prior GPT-6 Sol cached rate (Sep 2026)
- Anthropic, Claude Sonnet 5.5 announcement (Sep 28, 2026)
- OpenAI, prompt-caching guide, cache-write billing and 30-minute lifetime (read Oct 8, 2026)
- ITdaily, Claude Sonnet 5.5 launch coverage (Sep 29, 2026)
