GPT-6.1 Sol vs Claude Sonnet 5.5: Same List Price, Different Agent Bills. Price the Workload Shape

Both list at $2/$10 with $0.10 cache reads, yet a 50-turn agent loop past 272K costs 29% more on one. Price your workload shape with this five-row worksheet.

Two identical price tags reading $2 input and $10 output per million tokens sit above a context ruler running past 300K tokens, with a step marked at 272K where one bill line jumpsTwo identical price tags reading $2 input and $10 output per million tokens sit above a context ruler running past 300K tokens, with a step marked at 272K where one bill line jumps
Same sticker on both tags; the step at 272K is where GPT-6.1 Sol vs Claude Sonnet 5.5 stops being a tie.

Run a 50-turn agent loop at a steady 100K tokens of context and the two price sheets agree to the cent: $2.10 on GPT-6.1 Sol, $2.10 on Claude Sonnet 5.5. Let the same loop grow by 5K tokens a turn until it ends at 345K and the bills split, $3.51 against $2.71 (illustrative, at list prices). Nothing about the per-token sticker changed between those two runs. The shape of the work did.

That is the whole GPT-6.1 Sol vs Claude Sonnet 5.5 pricing story for anyone running agents rather than chats. Both models list at $2 input and $10 output per million tokens, both charge $0.10 for a cache read, and one of them reprices an entire request once its input crosses 272K tokens. A worksheet that prices your actual loop shapes on both sheets tells you which lanes are exposed. A sticker comparison tells you nothing.

This piece gives you that worksheet with five shapes filled in, the arithmetic behind each row, an alarm to set before the long-context line, and the recompute rule for the day either page changes. It does not pick a model. Model quality, tokenizer efficiency and task success sit outside a price sheet, and cost per completed task is still the number that settles a lane.

The two price sheets as read on Oct 7: $2/$10, $0.10 reads and one 272K line

Anthropic released Claude Sonnet 5.5 on September 28, 2026, per its announcement, with launch coverage such as ITdaily’s the next day. OpenAI followed on September 29 at DevDay with GPT-6.1 Sol, per DataCamp’s launch write-up; OpenAI’s launch post reportedly carries the same date but refused our fetch. The API model is gpt-6.1-sol. It is not the “GPT-6 Sol” that ChatGPT now uses as its chat default, and prices for one say nothing about the other.

We read both pricing pages on the evening of October 7, US time (about 00:20 to 00:45 UTC on October 8), and re-read them before filing. OpenAI’s pricing page lists gpt-6.1-sol at $2.00 input, $0.10 cached input, $2.50 cache writes and $10.00 output per million tokens for short context. Its long-context columns for the same row read $4.00, $0.20, $5.00 and $15.00.

The page defines the line exactly: “Short context: ≤272K input tokens. Long context: >272K input tokens.” The GPT-6.1 Sol model page adds that the long-context rate applies to the full request, and lists a 1,050,000-token context window with 922,000 maximum input tokens and 128,000 output.

OpenAI Developers pricing page showing the Flagship models table with gpt-6.1-sol at $2.00 input, $0.10 cached input, $2.50 cache writes and $10.00 output for short context, and $4.00 and $0.20 in the long context columns Screenshot: OpenAI Developers, “Pricing | OpenAI API” (flagship models table, Oct 7, 2026), captured Oct 7, 2026.

Anthropic’s pricing documentation lists Claude Sonnet 5.5 at $2 input, $2.50 for a 5-minute cache write, $4 for a 1-hour cache write, $0.10 for cache hits and refreshes, and $10 output per million tokens. Its footnote prices Sonnet 5.5 cache hits at 0.05x base input. The consumer-facing claude.com pricing page shows the same $2, $10, $2.50 and $0.10.

Claude Platform Docs pricing page with the model pricing table showing Claude Sonnet 5.5 at $2 per MTok input, $10 output, $2.50 for 5-minute writes, $4 for 1-hour writes and $0.10 cache hits Screenshot: Claude Platform Docs, “Pricing - Claude Platform Docs” (model pricing table, Oct 7, 2026), captured Oct 7, 2026.

The long-context section is where the sheets part ways. Anthropic says Claude 4.6 and later models, except Haiku 5.5, “include the full 1M token context window at standard pricing,” and spells it out: “A 900k-token request is billed at the same per-token rate as a 9k-token request.” Batch is 50% off on both vendors, and Anthropic’s caching multipliers stack with its batch discount.

Claude Platform Docs long context pricing section stating that a 900k-token request is billed at the same per-token rate as a 9k-token request, and that Claude Haiku 5.5 is priced by prompt length Screenshot: Claude Platform Docs, “Pricing - Claude Platform Docs” (long context pricing passage, Oct 7, 2026), captured Oct 7, 2026.

Three corrections belong next to those prices. First, if you have seen "cache reads $0.10 vs $0.20" in a comparison, it is wrong as of today. Both models list $0.10. The $0.20 figure is GPT-6.1 Sol's cached rate past 272K, GPT-6 Sol's old cached rate per DataCamp, and the cache-read price legacy Sonnet 5 still lists. It was also Sonnet 5.5's own rate until Anthropic's [Claude Haiku 5.5 announcement](https://www.anthropic.com/claude-haiku-5-5) on October 7 halved Sonnet 5.5 cache reads to $0.10, nine days after launch.

Second, OpenAI’s table shows a $2.50 cache-writes column for gpt-6.1-sol, and its prompt-caching guide says caching is on by default and that cache writes on GPT-5.6 and later cost 1.25× the uncached input rate. Our worksheet charges every new token as a write. Tokens placed after the last explicit cache breakpoint bill at plain input instead; if every new token billed that way, the flat-100K loop would cost $1.975 instead of $2.10. Third, the same guide keeps a cached prefix for 30 minutes after its last write or reuse, against Anthropic’s five-minute default, so the pause scenario below is priced for Anthropic only.

Why an agent loop exposes what one chat request hides

A single chat turn at 20K tokens bills about the same anywhere with a $2/$10 sticker. An agent loop is different in two ways that the sticker hides. It re-sends its whole history every turn, so most of what you pay for is cache reads, and its history grows, so a lane that starts at 100K can finish well past 272K without anyone deciding it should. Both effects compound per turn, and both are invisible until the invoice arrives.

Agents also pause. A lane waiting five minutes for a human approval can come back to an expired cache and pay to write its entire prefix again. A one-vendor cache meter catches that inside a single sheet; the worksheet below catches it across two.

Build the workload-shape worksheet in six steps

Step 1: Log four numbers per turn for every lane

You cannot price a shape you have not measured. For each agent lane, record per turn the input tokens sent, how many of those were cache reads, how many were cache writes, and the output tokens, reasoning included. Add the wall-clock gap since the previous turn, because that gap decides whether the cache survived. One line per turn is enough; the field names below are illustrative.

{"lane": "repo-refactor", "vendor": "openai", "model": "gpt-6.1-sol",
 "turn": 36, "input_tokens": 275000, "cache_read": 270000, "cache_write": 5000,
 "output_tokens": 2000, "gap_since_prev_s": 41, "tier_billed": "long",
 "sheet_read": "2026-10-07"}

Measure on both vendors when a lane runs on both. Anthropic’s own pricing page notes that its current tokenizer “produces approximately 30% more tokens for the same text” than its previous one, so the same transcript is not the same token count across tokenizers. The worksheet below assumes equal counts to isolate the price mechanics. Your log should replace that assumption with measured counts.

Step 2: Sort lanes into five workload shapes

Most agent lanes fall into one of five shapes. Name the shape in the lane’s config so the bill has a reason attached.

Classify from a week of logs, not from the lane’s intent. A lane described as “quick triage” that keeps every tool result in history is a growing lane, whatever its owner calls it. Plot input tokens per turn for the ten longest runs; a flat line, a ramp and a step at load time are the three patterns that matter, and output share tells you whether the cache price even moves the total.

Shape What it looks like Typical lane
Flat 100K × 50 turns Context stays near 100K; old turns get compacted Support triage with a fixed knowledge pack
Growing 100K → 345K × 50 turns Each turn adds about 5K and nothing is dropped Long refactor or research loop
Flat 300K × 50 turns A large repo or document set loaded up front Codebase-wide review
Mostly-cached 20K × 50 turns Small context, output-heavy turns Drafting, form filling, short tool calls
Batch overnight, 100 × 300K Independent requests, no cache, results by morning Document intake or bulk classification

Step 3: Price every shape on both sheets

This is the artifact. Every number is illustrative: list prices as read on October 7 (US evening), standard tier unless the row says batch, no data-residency uplift, equal token counts on both vendors, and every new token per turn billed as a 5-minute cache write. The re-run column is the date you owe the next recompute.

Workload shape Input tokens (run) Cached share Output tokens GPT-6.1 Sol bill Sonnet 5.5 bill Delta Mechanic that caused it Re-run date
Flat 100K × 50 turns 5,000,000 95% 100,000 $2.10 $2.10 $0.00 Both under 272K; identical rates Next change to either pricing page
Growing 100K → 345K × 50 turns 11,125,000 97.8% 100,000 $3.51 $2.71 +$0.80 (+29%) Turns 36–50 pass 272K; OpenAI reprices the whole request Next change to either pricing page
Flat 300K × 50 turns 15,000,000 98.3% 100,000 $5.70 $3.10 +$2.60 (+84%) Every turn past 272K: 2× reads and writes, 1.5× output Next change to either pricing page
Mostly-cached 20K × 50 turns 1,000,000 95% 50,000 $0.72 $0.72 $0.00 Under the line; output is 69% of the bill Next change to either pricing page
Batch overnight, 100 × 300K 30,000,000 0% 200,000 $61.50 $31.00 +$30.50 (+98%) Batch halves both sheets; the 272K line still applies on OpenAI Next change to either pricing page

The arithmetic for the rows that differ, so you can check it against your own reading of the pages:

Growing 100K -> 345K, 50 turns, 5K new + 2K out per turn (illustrative)
  context at turn k = 100K + 5K(k-1); turn 35 = 270K, turn 36 = 275K
  cache reads = 50 x 95K + 5K x 1,225 = 10,875,000; writes 250,000; output 100,000
  Sonnet 5.5:      10,875,000 x 0.10 + 250,000 x 2.50 + 100,000 x 10 = 2.7125
  GPT-6.1 Sol 1-35: 6,300,000 x 0.10 + 175,000 x 2.50 +  70,000 x 10 = 1.7675
  GPT-6.1 Sol 36-50: 4,575,000 x 0.20 + 75,000 x 5.00 +  30,000 x 15 = 1.7400
  GPT-6.1 Sol total = 3.5075 (+29%)      (all prices USD per million tokens)

Flat 300K, 50 turns: 295K read + 5K write + 2K out per turn
  Sonnet 5.5:  0.0295 + 0.0125 + 0.020 = 0.062 a turn -> 3.10
  GPT-6.1 Sol: 0.0590 + 0.0250 + 0.030 = 0.114 a turn -> 5.70

Batch overnight, 100 requests x 300K input, 2K out, no cache
  Sonnet 5.5 batch ($1 in / $5 out):            0.30 + 0.010 = 0.310 -> 31.00
  GPT-6.1 Sol batch long ($2.00 in / $7.50 out): 0.60 + 0.015 = 0.615 -> 61.50

In the growing loop, the whole $0.80 gap sits in the last 15 turns. Sonnet 5.5 bills those turns $0.945 and GPT-6.1 Sol bills them $1.74. Turns 1 through 35 cost the same to the cent on both. The batch row uses the long-context batch rates OpenAI lists for this model ($2.00 input and $7.50 output once past 272K), which is why a 50% discount does not close the gap.

Read the two zero-delta rows as carefully as the others. The mostly-cached 20K lane spends 69 cents of every dollar on output, so the cache-read price barely moves it and the 272K line never comes near. Moving that lane between vendors on price grounds would be a waste of a migration. The flat-100K lane is equally safe until somebody removes its compaction step, at which point it becomes the growing shape without changing its name.

The batch row is the one most teams miss. Batch jobs feel cheap by construction, and nobody watches their per-request size because nobody watches them at all. Take a document-intake job whose inputs creep from 250K to 300K over a quarter: it crosses the line quietly, and its bill more than doubles on one vendor while its request count stays flat.

Grouped horizontal bar chart, illustrative, comparing GPT-6.1 Sol and Claude Sonnet 5.5 bills for three 50-turn shapes: flat 100K both $2.10, growing to 345K $3.51 versus $2.71, flat 300K $5.70 versus $3.10Grouped horizontal bar chart, illustrative, comparing GPT-6.1 Sol and Claude Sonnet 5.5 bills for three 50-turn shapes: flat 100K both $2.10, growing to 345K $3.51 versus $2.71, flat 300K $5.70 versus $3.10 Illustrative: three 50-turn loop shapes priced from both list-price pages read Oct 7, 2026; equal token counts assumed.

Step 4: Put a context alarm at 80% of the vendor’s line

The cliff is a property of the vendor, not of your task, so the alarm belongs in the lane’s config next to the model ID. For a 272K line, 80% is about 218K. When a lane’s 95th-percentile input per turn crosses that, compact or summarize before the next turn, not after the bill shows it. A lane on a vendor with no long-context tier gets no alarm for this reason, though it may still want one for latency.

lane: repo-refactor
model: gpt-6.1-sol
price_sheet_read: 2026-10-07
long_context_line_tokens: 272000
alarm_at_tokens: 218000      # 80% of the line
on_alarm: compact_history    # summarize oldest turns before the next request
p95_window_turns: 20

Flow diagram: an agent turn request goes to a context-size count, past an alarm at 218K, into a cliff check at 272K that routes to the standard tier or the long-context tier, both ending on the bill lineFlow diagram: an agent turn request goes to a context-size count, past an alarm at 218K, into a cliff check at 272K that routes to the standard tier or the long-context tier, both ending on the bill line Every request passes the same check; the alarm at 218K gives the lane one chance to compact before the whole request reprices.

Compaction is not free either. A summary turn costs output tokens and can drop facts the agent later needs. Price it in the worksheet as its own row if a lane compacts often.

Step 5: Price the pauses against the cache lifetime

Anthropic’s 5-minute cache write lasts five minutes; the 1-hour write costs $4 per million tokens and lasts an hour. An agent that waits longer than five minutes for an approval comes back to a cold cache. In the flat-100K shape, that turn re-writes its whole 100K prefix at $2.50 per million, so it costs $0.25 plus $0.02 of output: $0.27 against the usual $0.042, about 6.4 times the turn price (illustrative).

A lane that pauses at every step should either buy the 1-hour write or keep its human gates short. Log approval-wait minutes beside each lane’s cache lifetime, and the worksheet will show you which one you are paying for.

OpenAI’s prompt-caching guide gives gpt-6.1-sol a 30-minute cache lifetime, renewed by each write or reuse, so the same pause costs nothing extra on that lane until it passes half an hour. Measure it anyway: send the same prefix after a known gap and check whether the response reports cached tokens.

Step 6: Recompute the worksheet when either page changes

Sonnet 5.5’s cache-read price halved nine days after launch. A worksheet filled in on September 30 had the wrong number in every Sonnet 5.5 row by October 7. Treat both pricing pages as inputs with a read date, not constants.

The recompute takes ten minutes if the worksheet keeps its inputs separate from its results. Store the eight gpt-6.1-sol cells and the five Sonnet 5.5 cells in one small table with the read date and URL, and let every row reference them. When a page changes, you update one table and every bill recalculates, including the deltas you showed someone last week.

The rule set, short enough to pin above the worksheet:

  1. Price the shape, not the sticker. Every lane has a named shape and a worksheet row; nobody compares vendors on the headline rate alone.
  2. Put a context alarm at 80% of each vendor’s long-context line, in the lane’s config.
  3. Log cache-hit share per lane per day; a drop means pauses, compaction or a prompt change is re-buying the cache.
  4. Recompute the worksheet whenever either pricing page changes, and stamp the read date on it.
  5. Keep promotional and introductory prices out of this worksheet; their expiry dates live in a separate price-expiry register.

If a lane bills in credits or a subscription allowance rather than API tokens, it does not belong here at all; translate it first with the meter dialects guide, and keep plan limits on the subscription plans side.

GPT-6.1 Sol vs Sonnet 5.5 bill failures, with the signal and the first fix

What breaks The signal you would see First action
A growing lane drifts past 272K on GPT-6.1 Sol Per-turn cost steps up mid-run while output per turn stays flat; usage shows long-context billing Turn on the 218K alarm and compact; re-price the lane as the growing shape
A comparison deck still uses “$0.10 vs $0.20” Sonnet 5.5 rows priced at twice the cache-read rate on today’s page Re-read both pages, correct the rows, stamp the read date
Approval pauses re-buy the cache Cache writes spike on the first turn after each human gate; cache-hit share falls Log wait minutes per gate; test the 1-hour write on that lane
Equal token counts assumed across vendors Measured tokens per task differ from the worksheet’s input column Replace modeled counts with logged counts from Step 1
A batch job sized under the line grows over it Overnight batch bill roughly doubles with no change in request count Check document sizes against 272K; split or chunk inputs before submission
OpenAI cache-write billing differs from the worksheet Invoice lower or higher than the worksheet for flat lanes Reconcile one day of usage against both readings; record which one the invoice matches

Keep one worksheet for every model lane in the fleet

A fleet with lanes on two vendors has two price sheets, two context lines and two cache lifetimes moving underneath it. The worksheet belongs where the lanes are assigned, not in a spreadsheet one person maintains. When the mode-to-model assignment table moves a lane from one vendor to another, the worksheet row moves with it and gets re-priced on the new sheet the same day.

That is the job a multi-agent command center does for the fleet: one place that shows each lane’s model, its context trend against the vendor’s line, and the read date of the price sheet it was budgeted on. The sticker is the same on both vendors this week. The bills are whatever your shapes make them.

FAQ

Is GPT-6.1 Sol cheaper than Claude Sonnet 5.5?

On list price they tie: both are $2 input and $10 output per million tokens with $0.10 cache reads. They diverge on long context. Past 272K input tokens OpenAI bills the whole request at 2× input and cache rates and 1.5× output, while Sonnet 5.5 stays flat across its 1M window.

Do GPT-6.1 Sol and Claude Sonnet 5.5 charge the same for cache reads?

Yes, as of October 7, 2026, both list $0.10 per million cached tokens. Sonnet 5.5’s rate was halved from $0.20 that day in Anthropic’s Haiku 5.5 announcement. GPT-6.1 Sol’s cached rate rises to $0.20 once a request passes 272K input tokens. Re-read both pages before budgeting.

How do I estimate what an agent loop costs on each model?

Log input, cached, written and output tokens per turn, then price each turn on both sheets, applying OpenAI’s long-context rates to any request over 272K. Sum the turns. A 50-turn loop at 100K costs $2.10 on both; the same loop growing to 345K costs $3.51 versus $2.71 (illustrative).

Sources

YOU'RE THROUGH THIS ONE.

Keep connecting the dots.

Back to the library