Token Plans Decoded: What Heavy AI Users Actually Pay in 2026

AI subscription plans decoded for heavy users: Claude Max, ChatGPT Pro, Copilot's new AI credits, Chinese flat plans, and the math that picks your stack.

Token plans decoded — how AI subscription plans sell smoothed capacity, not tokens
Every plan is a capacity bet; the caps mark where the bet stops covering you.

On June 1, 2026, GitHub Copilot stopped counting requests and started metering tokens. A year before that, Anthropic stacked weekly caps on top of its five-hour usage windows. OpenAI sells credit top-ups for the day your Codex allowance runs dry. None of this is random. It is the same correction arriving vendor by vendor: AI subscription plans were priced for the median user, and if you run agentic CLIs hard enough to hit caps weekly, you are not the median user.

Start with the reframe that decodes everything below: no vendor sells tokens — they sell smoothed capacity. A flat monthly price is a bet that your average consumption costs less to serve than your subscription. Every confusing mechanic on every pricing page — rolling windows, layered caps, model multipliers, credit grants — exists to manage the gap between the subscriber that bet was priced for and you.

This piece decodes the whole market for the heavy user: what Claude Max, ChatGPT Pro, Google’s AI tiers, the newly metered GitHub Copilot, and the Chinese flat plans actually sustain, and the arithmetic for choosing between upgrading, stacking, and API overflow. Three payment models and six mechanics cover every plan on sale. Learn them once and any new pricing page reads in thirty seconds.

One honesty flag before the numbers: this market reprices monthly. Every figure below was checked on August 27, 2026 and will drift; the mechanics outlive the prices, which is why the mechanics come first.

The three ways AI capacity is sold

Flat subscription with usage windows. You pay monthly; the vendor books predictable revenue and bets on average usage. The caps are not spite — they are the precise line where the cross-subsidy ends. Median chat users fund the margin; the P99 agent user consumes it; windows and caps keep the second group from inverting the economics. This model wins for you when your usage is steady, near-daily, and lives mostly inside one vendor’s models.

API pay-as-you-go. True metering: every input and output token is priced, cost tracks consumption exactly, and nobody subsidizes anybody. It feels expensive precisely because it is honest. Open-weight competition has dragged open-model API list prices down roughly 80 percent year over year, per mid-2026 reporting, and the DeepSeek effect on agent economics is the reason a hard price floor now sits under the entire market. This model wins for you when usage is spiky, attributable to projects or clients, or routed across many models.

Hybrid: a subscription floor with a metered ceiling. A flat base plan plus credits or overage once included capacity runs out — Codex credits, Copilot’s AI-credit top-ups, Anthropic’s buy-extra-usage option at cap, Devin’s twenty-dollars-plus-usage repricing. Hybrids exist to capture heavy-user upside without unlimited vendor liability. This model wins for you when you want plan pricing most days and burst capacity on the bad ones.

Read the past fourteen months of plan redesigns together and one direction appears: everything moved toward the hybrid. Copilot’s June switch is the loudest example, and the pattern predicts where every remaining flat plan’s caps are heading before any announcement does.

Six mechanics that decide what an AI subscription plan actually delivers

The vendor sections below reference these six mechanics instead of re-explaining them each time. Verified as of August 27, 2026.

Mechanic What it means Where you’ll meet it
Rolling window Quota measured over a sliding period — commonly 5 hours — starting from your first message, not the calendar day Claude plans; ChatGPT and Codex limits
Layered caps A weekly or monthly ceiling stacked on the rolling window; bursty users hit the window, steady users hit the ceiling Claude weekly caps; Codex weekly limits
Model multipliers / credit rates Identical requests debit different amounts depending on model class Copilot AI credits; Fable-class vs Sonnet-class burn
Degradation and fallback What happens at the cap: silent fallback to a smaller model, queue deprioritization, or a hard stop Antigravity CLI’s free-lane fallback; tier-based routing
Overflow purchase Buying extra metered capacity at API-adjacent rates once the plan runs out Claude extra usage; Codex credits; Copilot top-ups
Tier gating Frontier models or reasoning modes reserved for upper tiers regardless of remaining quota Fable 5 access by tier; Pro-only reasoning modes

Six mechanics of AI subscription plans: windows, caps, multipliers, fallback, overflow, gating The six levers every vendor pulls; only the settings differ.

Two of the six do most of the work. The window-plus-cap pair determines whether your usage shape fits a plan at all, and the multiplier mechanic determines how fast your model mix drains it. Check those two on any pricing page and you have 80 percent of the decision; the rest is fine print about what failure feels like at the limit.

Anthropic: Claude Pro, Max 5x, and Max 20x

Anthropic’s lineup is the cleanest expression of the flat-plan-with-windows model, and the Claude plan documentation is unusually explicit about the mechanics. Verified as of August 27, 2026.

Plan Price Mechanics What it sustains in Claude Code
Pro ~$20/mo 5-hour rolling window + weekly cap A few focused hours of Sonnet-class work daily; limited Claude Fable 5 access
Max 5x ~$100/mo Same structure at ~5x Pro’s usage The single-session daily driver: a full workday, most days
Max 20x ~$200/mo Same structure at ~20x Pro’s usage Parallel sessions and a heavy Fable 5 habit

The naming is literal. Max tiers are usage multiples of Pro inside the same rolling-window-plus-weekly-cap structure — same models, same surfaces, nothing else changes — so the only real question is which multiple your working week needs.

Community-measured reports, not vendor promises, are the right calibration source, and they consistently sketch the same shape: Pro sustains a few focused hours of Claude Sonnet 5 work per day before the window bites; Max 5x holds up as a full-workday single-session tier; Max 20x is for people running parallel sessions or leaning on Claude Fable 5 for most steps.

Two pressure valves before you upgrade. Model mix is the biggest burn lever — Fable-class tokens debit far more quota than Sonnet-class, so routing routine steps down-tier stretches any plan. And Max tiers can buy extra usage at API-adjacent rates when a cap hits, which quietly turns the plan into a hybrid on your worst days. Our Claude Code field guide covers the craft of stretching a Claude budget in detail.

OpenAI: ChatGPT Plus, Pro, and Business — and where Codex fits

OpenAI bundles its agent product into the plans most developers already pay for, which makes the tier question mostly a Codex question. Verified as of August 27, 2026; mechanics are documented across OpenAI’s platform docs.

Plan Price What it covers for agent work
Plus ~$20/mo Codex included with modest rolling-window limits — fine for evening sessions, tight for daily driving
Pro ~$200/mo The heavy-Codex tier: much larger windows, priority capacity, frontier reasoning modes
Business ~$25–30/seat/mo Plus-class limits per seat, workspace controls, training excluded by default

Codex usage debits against your plan through the familiar window-plus-weekly-cap pattern, and when included capacity runs out OpenAI sells credits — metered top-ups that keep a session going at pay-per-use rates rather than forcing an upgrade. That credit lane is the hybrid model in its purest form. Our OpenAI Codex review works the plan math surface by surface, including where cloud tasks bill differently from CLI turns.

Is ChatGPT Pro worth $200 a month? Worth it if you run Codex daily as a delegation engine, fan out parallel cloud tasks, or hold long frontier-reasoning sessions — Pro’s windows are sized for exactly that, and the people doing it report the cap effectively disappears from their week. Skip it if Plus covers your window most days and an occasional credit top-up or API call absorbs the spikes; that combination usually lands well under a third of Pro’s price.

Google: AI Pro, AI Ultra, and the free lane that died

Google’s paid tiers are simple enough — AI Pro near $20 a month, AI Ultra near $250 — bundling Gemini 3.1 app limits, coding quotas, and consumer extras, with developer surfaces documented at Google’s AI developer site. The interesting story is what happened to the free lane.

For a year, the category’s famous free lunch was Gemini CLI on a personal Google account: request quotas generous enough that heavy users routed grunt work there and saved paid windows for judgment calls. That lane is gone. Google discontinued Gemini CLI on June 18, 2026 — a shutdown abrupt enough to break CI pipelines that had scripted against it — and pointed everyone at its Go-based successor, Antigravity CLI. The full shutdown-and-migration story is in our Gemini CLI and Antigravity guide.

Free lane, 2026 status: Antigravity CLI ships its own personal-account free quota, tighter than Gemini CLI’s was, with fallback to Flash-class models at the limit. Treat any preview-labeled quota as changeable without ceremony — June 18 was the proof.

Ultra decodes as a bundle, not a bigger coding plan: its value rides on video models, storage, and consumer extras rather than agent capacity. Most readers of this piece want AI Pro plus a cheap volume lane, and should let Ultra sell itself to someone else.

GitHub Copilot: premium requests are out, AI credits are in

GitHub Copilot AI credits are the usage meter that replaced premium requests on June 1, 2026. Every paid Copilot plan includes a monthly credit grant; chat, agent-mode, and premium-model usage debit credits in proportion to the tokens consumed, and heavier models drain the grant faster. Extra credits are metered purchases.

The old system deserves one paragraph because the internet is still full of it. From mid-2025, Copilot plans carried monthly allowances of “premium requests” — roughly 300 on the $10 Pro tier and 1,500 on the $39 Pro+ — debited by per-model multipliers, with overage around $0.04 a request. Forum threads decoding multiplier tables became a genre of their own. All of that is dead terminology now, and any guide still explaining multipliers is describing the previous era.

GitHub announced the change in spring 2026 and framed it as moving Copilot to usage-based billing: agent-mode sessions vary too wildly in cost for flat request-counting to price honestly. Developers read the same fine print differently — as Visual Studio Magazine’s coverage of the reaction put it, “you will get less, but pay the same price”. Both readings are correct. Token metering prices agent work more honestly, and honest pricing is worse pricing for the heavy users who were being cross-subsidized.

Here is what the switch means inside one session. An agent-mode run that reads a large repo and grinds tests debits credits for every token it touches — a single long session can now consume what a week of old-style requests did, while light autocomplete-and-chat use barely dents the grant. Org-level budget caps and overage controls exist; set them before your first agent-heavy sprint, not after the invoice. Copilot’s move is the clearest signal yet that the flat-rate era is wobbling, and we cover the market-wide pattern in the subscription squeeze.

The rest of the US field, quickly. xAI’s SuperGrok tiers (~$30 and ~$300) buy Grok capacity with agent features attached — sensible if Grok is already your model, niche otherwise. Amazon Q Developer still lists a free tier and a ~$19 Pro seat, but new signups have reportedly been blocked since mid-2026, which makes it a lane for existing AWS shops only. And Cognition’s Devin — whose desktop IDE is the former Windsurf, now Devin Desktop — repriced from its famous $500 tier to $20 a month plus usage: the hybrid model again, arriving from the opposite direction.

The Chinese undercut: flat plans priced as substitutes

The most aggressive pricing in the market comes from Chinese vendors selling flat coding plans as direct substitutes for Claude and ChatGPT tiers — and as of August 2026 the undercut has not blinked. Verified as of August 27, 2026.

Plan Entry price What it is
Z.ai GLM Coding Plan ~$3/mo entry tier GLM-5.2 quota marketed explicitly against Claude plans on price-per-usage multiples
Kimi For Coding ~$19/mo Kimi K3 capacity packaged for coding harnesses
MiniMax coding plan ~$10/mo The third flat-plan substitute, running the same playbook
DeepSeek (no plan needed) $0.14 / $0.28 per M tokens Raw API pricing so low it stands in for a plan

Z.ai anchors the segment: the GLM Coding Plan’s entry tier sits near $3 a month, and its marketing leans on usage-multiple comparisons against Claude tiers rather than benchmarks. DeepSeek skips plans entirely because DeepSeek V4 Flash’s API pricing — $0.14 per million input tokens, $0.28 out — is the credible price floor for agentic coding. At that rate, metering is the plan.

One free lane closed here too. Qwen Code’s free OAuth tier — for a while the most generous free allowance in the category — ended April 15, 2026, per the mid-2026 harness map; Qwen work now routes through paid Alibaba Cloud API keys at rates that are still cheap but no longer zero.

The substitution mechanics are the actual story. These plans expose Anthropic- and OpenAI-compatible endpoints, so they plug into harnesses you already run — often a base-URL swap in a config file, no new tooling. Our Chinese CLI wave review covers the tool side. Jurisdiction and data terms differ sharply by provider and routing path; our Chinese frontier model guide carries the trust framework, and this piece won’t re-litigate it.

The heavy-user math: what agent workloads actually burn

Agents hit caps chat never touches because of how the loop works: every step re-sends the growing context, so a two-hour session doesn’t consume tokens linearly — it compounds. Input tokens dominate output by an order of magnitude, and long sessions multiply quietly while you watch the diff scroll. How harnesses drive this burn is its own article; the budget consequence is what matters here.

Ground it in one week. Take an honestly heavy week from a metered log : a two-day refactor on a legacy service, a test-backfill grind, and two evenings running two parallel sessions. Call it 40 agent-hours. Weeks shaped like that land in the hundreds of millions of input tokens once every loop iteration is counted — which is why one hard week can eat a disproportionate slice of any weekly cap while an ordinary week never gets close.

Now price that same week under three setups:

A. One Max-class plan B. Mid plan + API overflow C. Mid plan + cheap stack
Setup Claude Max 20x Claude Max 5x + Anthropic API key Claude Max 5x + GLM Coding Plan + DeepSeek V4 Flash API
Monthly cost ~$200 flat ~$100 + ~$40–90 metered ~$110–125 all-in
What ran out first Weekly frontier cap, Thursday afternoon The 5-hour window, daily Nothing hard-stopped — attention did
Where the week hurt Parallel evenings throttled; Friday demoted to Sonnet-class Refactor days spiked the overflow bill Three dashboards, three session archives

Three-scenario cost comparison of one heavy agent week across AI subscription setups One heavy week, three ways to pay for it. The shape of your usage picks the winner.

The table earns one finding: usage shape picks the winner, not sticker price. Steady all-day single-session work favors A, because a flat cap you rarely hit is the cheapest capacity sold. Bursty weeks with quiet stretches favor B, because overflow only bills when you spike. Parallel fleets and volume grinding favor C, because cheap lanes absorb exactly the steps that drain frontier caps fastest.

Stacking AI subscription plans: the meta, handled honestly

What power users actually do, observed rather than recommended: hold two or three plans across vendors, route each task to the cheapest adequate lane, and rotate toward whichever window has reset. Running multiple coding agents without the chaos systematizes exactly this posture.

Draw the terms line precisely, because there is one. Holding plans with multiple vendors is unambiguously fine — nobody’s terms restrict who else you pay. Running multiple accounts on one vendor to evade limits, or buying resold account access, generally violates terms of service and gets accounts banned. Stay inside the lines: stack vendors, never accounts. The economics work without cheating.

The routing heuristic that makes a stack pay: frontier plan for judgment-heavy steps — planning, architecture, review — and cheap lanes for volume steps — boilerplate, summaries, test loops. It is the model-routing pattern from the DeepSeek economics piece applied one level up, to plans instead of API calls.

And name the overhead honestly: every added plan is another dashboard, another session archive, another place usage hides. Stacking without measurement is how people pay $150 a month to feel clever when a measured $100 setup would have covered them. The last two sections of this piece are the antidote.

API vs subscription: the crossover math

A plan is prepaid capacity at an implied discount. Price its included usage at API list rates and you know exactly what walking away forfeits — and community measurements consistently put heavy-user Max-class consumption at several multiples of sticker price when valued at API rates, which is the entire argument for plans under steady heavy usage.

Three rules cover the decision:

Usage shape Best structure Why
Light or spiky API pay-as-you-go You would forfeit unused plan capacity most weeks
Steady and heavy Flat plan, or two The cross-subsidy finally runs in your favor
Extreme or parallel Hybrid: plan + overflow, or a stack Caps bind before the plan’s value runs out

API-only wins regardless of shape in four cases: variable team workloads where seats would idle, gateway routing across many models, per-project cost attribution for client billing, and workflows built on models no plan covers — which in August 2026 mostly means the open-weight lane, where the same crossover math runs against much smaller numbers.

Track your actuals: the discipline under every plan choice

The uncomfortable premise: most people reading this cannot say what last week burned per provider. Every upgrade, downgrade, and stacking decision made in that state is a guess wearing a spreadsheet costume.

Track four things: tokens in and out per provider and model, cap-hit events, where in the week those hits land, and your model mix. Those are precisely the inputs the crossover math needs — nothing more exotic than that.

Vendor dashboards will not do this for you. They are per-vendor silos, token detail lags or is absent entirely, and nothing aggregates across providers. Measurement has to live where the sessions live: locally, beside the harnesses.

Then make it a ritual. On renewal day, review measured burn against each plan’s implied budget, and downgrade, stack, or add overflow based on the delta. Ten minutes a month turns every number in this article from interesting into actionable.

Product note: Automater Lite meters token usage locally across every provider you run — Claude Code, Codex, Antigravity CLI, and 10+ more — so renewal day is a data decision, not a guess. The opt-in leaderboard turns the same numbers into proof-of-work. Free on automater.ai.

Automater Lite Usage panel showing weekly usage indicators for Claude, Codex, Kimi and Grok.
Automater Lite puts usage from several agent providers side by side. Values shown are a captured example, not plan allowances. Source: Automater · License and attribution.

Where plan pricing is heading

The loss-leader question first, handled carefully: reporting and vendor behavior through 2025 and 2026 suggest flagship plans lose money on their heaviest users, and the caps, multipliers, and credit systems appearing at every vendor are the visible correction rather than coincidence. Treat the claim as reported; treat the mechanics as evidence.

Predictions, stated so they can be wrong. Entry prices hold near $20, because Chinese flat plans and open-weight APIs put a hard floor under what anyone can charge — falsified if sticker prices broadly rise instead. Caps tighten and overflow monetization spreads to every remaining flat plan — falsified if a major vendor removes a cap class. Model gating by tier deepens as frontier releases get more expensive to serve — falsified if frontier access flattens across tiers instead.

Watch signals worth a calendar reminder: new cap classes appearing mid-cycle, credit systems arriving at vendors that lack them, and tier-gated model launches. Each one is a vendor telling you where its heavy-user losses live.

The close is agency. Re-run the crossover math quarterly against your own measured burn. Every number in this piece will churn; the three payment models and six mechanics will still be the whole game.

FAQ: AI subscription plans in 2026

What is the Claude Max plan and what are its limits?

Claude Max is Anthropic’s heavy-usage tier in two versions — roughly $100 and $200 monthly — offering about 5x and 20x Pro’s usage under the same five-hour rolling window plus weekly cap structure. Max tiers can also buy extra usage at metered rates when caps hit.

Is ChatGPT Pro worth $200 a month?

Yes for daily heavy Codex use, parallel cloud delegation, and long frontier-reasoning sessions — Pro’s windows are sized for exactly that. No if Plus covers most days: occasional credit top-ups or API overflow on spike days usually totals well under Pro’s price.

What happened to GitHub Copilot premium requests?

GitHub retired premium requests on June 1, 2026 and replaced them with AI credits — usage-based billing where every chat, agent, and premium-model interaction debits credits in proportion to tokens consumed. Plans include monthly credit grants; heavy agent users burn them faster than the old request allowances.

Are AI subscriptions cheaper than the API?

For steady heavy use, yes — plans sell capacity below API list price, and heavy users extract multiples of sticker. For spiky, light, or attributable usage, the API wins because you pay only for consumption. Extreme users need both: plans for the base load, metering for spikes.

Can I stack multiple AI subscriptions?

Across vendors, yes — holding Claude, ChatGPT, and a Chinese flat plan simultaneously is common and violates nothing. What breaks terms is multiple accounts on one vendor to evade limits, or resold account sharing. Stack vendors, route tasks to the cheapest adequate lane, and stay inside terms.

What is the cheapest way to run coding agents?

Combine a cheap flat plan — the GLM Coding Plan’s entry tier is near $3 monthly — with rock-bottom open-weight APIs like DeepSeek V4 Flash at $0.14 per million input tokens, reserving any frontier plan for judgment-heavy steps. Free lanes still exist but shrank sharply through 2026.

Sources