The Chinese CLI Wave: Qwen Code, Kimi Code CLI, and the $3 Coding Plan

Qwen Code, Kimi Code CLI, and the Z.ai GLM Coding Plan reviewed for August 2026: real prices, quotas, Claude Code wiring recipes, and a calm trust checklist.

Qwen Code, Kimi Code CLI, and the GLM Coding Plan mapped as the 2026 Chinese CLI wave
The wave in one frame: four labs, three distribution strategies, one pattern.

The screenshots keep landing in your feeds: Claude Code mid-refactor, except the model behind it is GLM-5.2 on a plan that cost $3 the first month. Qwen Code grinding through a migration on tokens priced in cents. A Kimi K3 benchmark chart with a Claude model in second place. Each screenshot travels alone; almost nobody connects them as one story.

They are one story. Since early 2025, nearly every major Chinese lab has shipped an agentic CLI, a flat-rate coding plan, or both, and priced the bundle roughly an order of magnitude under the US plans you probably pay for today. This guide covers the wave as it actually stands in August 2026: Qwen Code after its famous free tier ended, Kimi Code CLI and the Kimi For Coding plan, the Z.ai GLM Coding Plan, what happened to iFlow CLI, and where DeepSeek fits despite shipping no flagship CLI of its own.

For each one you get the same honest treatment: what it is, what it costs, what data it touches, how to wire it into the harness you already use, and whether it belongs in your rotation. Prices, quotas, and endpoints were checked as of August 27, 2026; this corner of the market moves monthly, so treat every number as perishable.

The wave: what happened after DeepSeek

The causal chain starts with DeepSeek’s R1 moment in January 2025, when near-frontier reasoning at a fraction of frontier price stopped being hypothetical. We traced the economics in the DeepSeek effect on agent costs; the short version is that once capability-per-dollar collapsed, someone was going to build distribution on top of it. The CLIs and coding plans are that distribution layer, and they chase the coding-agent habit specifically, not chat.

By mid-2025 the wave had a shape, and by 2026 a name-brand roster: Alibaba’s Qwen Code, Moonshot AI’s Kimi tools, Zhipu’s GLM Coding Plan under the Z.ai brand, and a churn of smaller entrants. The strategy runs on two fronts at once. Open weights win the ecosystem — forks, hosting partners, benchmark presence, fine-tunes — while $3-to-$30 flat plans undercut Claude and ChatGPT subscriptions by enough that the price difference reads as a typo. Token plans decoded covers the plan economics across the whole market; this piece stays on the Chinese AI coding tools themselves.

One more frame before the reviews: none of these tools exists in a vacuum. They compete inside the broader agent harness field, where the 2026 pattern is brutal churn — Gemini CLI is gone, several 2025 darlings are gone, and the survivors won by being cheap, compatible, or both. The Chinese entries are aggressively both.

The roster, compressed — every claim in this table gets its full treatment below:

Vendor CLI Flat plan Entry price signal Endpoint compatibility
Alibaba Qwen Code (Apache-2.0) None — metered API $0.11/M input (Qwen3-Coder-Next) OpenAI-compatible
Moonshot AI Kimi Code CLI Kimi For Coding Flat tiers, under Claude plans Anthropic-compatible
Zhipu / Z.ai None — plan-first GLM Coding Plan (Lite/Pro/Max) ~$3/mo Lite promo Anthropic-compatible
DeepSeek No first-party flagship None — metered API $0.14/M input (V4 Flash) Anthropic-format + OpenAI-compatible
iFlow iFlow CLI Free aggregator Shut down mid-2026, per reporting

Qwen Code: the flagship after the free lunch

Qwen Code is Alibaba’s open-source agentic coding CLI: a fork of Google’s Gemini CLI, retuned — parser, prompts, and tool loop — for the Qwen3-Coder model family, and shipped under Apache-2.0 at github.com/QwenLM/qwen-code. It became the wave’s most-adopted piece for one famous reason: a free OAuth tier that offered on the order of 2,000 requests a day, no token counting, generous enough to daily-drive.

That reason is now history. The free OAuth tier ended April 15, 2026, per the mid-2026 harness map that has become the field’s informal census. Today you bring your own key — Alibaba Cloud’s international endpoints, ModelScope, or any OpenAI-compatible reseller — and pay per token. The consolation is that the tokens are cheap: Qwen3-Coder-Next, the current engine, lists at $0.11 per million input and $0.80 per million output on Alibaba Cloud, per the 2026 open-coding-model roundup. A week of heavy daily-driving on those rates typically lands in single-digit dollars, which softens the grief for the free tier considerably.

There is a second wrinkle worth savoring: Qwen Code has outlived its own upstream. Google discontinued Gemini CLI on June 18, 2026, which means the fork’s lineage is now maintained in one place — Alibaba’s. The Gemini inheritance still shows in the UX and tool loop, for better and worse, but divergence is permanent now.

The mini-review: best for high-volume implementation work on a metered budget, and for anyone who wants an open-source CLI whose engine also comes in a genuinely self-hostable size — Qwen3-Coder-Next runs on one ~46GB machine, with model-level detail in our Chinese frontier models survey. Standout: the cleanest weights-to-CLI story in the wave; one vendor covers editor, model, and license. Limits: quota and auth policy have now changed twice in eighteen months with modest notice, and the data-use terms on Alibaba’s hosted endpoints deserve a real read before production repos flow through. Cost: free software, metered tokens, no flat plan as of this writing.

Kimi Code CLI and Kimi For Coding: the tool-use bet

Moonshot AI runs the same two-piece playbook with different emphasis. The terminal agent shipped in 2025 as Kimi CLI, a technical preview with shell integration and MCP support; the current tool goes by Kimi Code CLI, per the mid-2026 harness map. The companion product, Kimi For Coding, is the flat plan — and, true to the wave’s pattern, it plugs into other harnesses as readily as into Moonshot’s own.

The bet underneath is the model line. Kimi K2 made the case in 2025 that a trillion-parameter-class MoE post-trained specifically for agentic tool use could out-behave bigger names inside a harness. Kimi K3, released July 16, 2026, escalated: a 2.8T-parameter MoE scoring roughly 93.4% on SWE-bench Verified in Vals AI’s independent testing, a 1M-token context window, and the first open-weight model to lead a frontend-coding arena ahead of Claude Fable 5. Tool-call reliability across long sequences is the design target, not a side effect, and it shows in harness work.

The plan itself is flat monthly pricing metered in requests per rolling window, positioned well under Claude’s subscriptions; Moonshot documents Anthropic-compatible endpoints so the plan can drive Claude Code directly, per Moonshot AI’s platform docs. We are deliberately not printing tier numbers here — Moonshot moved limits repeatedly during the preview era, and a stale price is worse than none.

Mini-review: best for long, tool-heavy agent runs where call-format discipline matters more than raw speed. Standout: the strongest independent benchmark story in the wave, and the clearest “trained to be an agent” positioning. Limits: the CLI remains younger and rougher than the model line, English docs trail the Chinese ones, and K3’s weights were staged to land after the API launch — check status before you plan on self-hosting. Cost: flat plan or metered API; the plan is the value play if your volume is real.

The Z.ai GLM Coding Plan: the $3 story

The Z.ai GLM Coding Plan is the wave’s clearest tell, because it is not a CLI at all. Zhipu’s international brand Z.ai sells a subscription whose entire pitch is to be the engine inside harnesses you already run — Claude Code first among them, with official setup guides. The plan comes in Lite, Pro, and Max tiers, with Lite marketed around $3 a month at promotional pricing, and it is the number that launched a thousand screenshots.

The tier structure, as listed at publish:

Tier Promo seen List price Quota shape
Lite ~$3/mo (intro) ~$6/mo ~120 prompts per 5-hour window
Pro ~$15/mo (intro) ~$30/mo ~600 prompts per 5-hour window
Max ~$30/mo (intro) ~$60/mo ~2,400 prompts per 5-hour window

Decode the marketing math before you buy. Tiers are sold as usage multiples of Claude’s plans and metered in prompts per five-hour window — precisely the budgeting structure Claude subscribers already know, which is not an accident. The multiples are vendor claims, not measurements; treat them the way you treat any benchmark a seller quotes about itself.

The engine grew teeth this summer. The GLM-4.5 and 4.6 era built the value reputation; the current model is GLM-5.2 — released June 16, 2026, MIT-licensed, 744B parameters with 40B active, and sitting at #1 among open models on the Artificial Analysis index, with ~78.7% on SWE-bench Verified in Epoch AI’s independent evaluation. Full model analysis lives in the Chinese frontier models survey.

The strategic read: a plan priced at pocket change, quota-framed like Claude’s, documented as a drop-in Claude Code substitute, backed by MIT weights anyone can host. If you want the whole wave’s strategy in one product, the Z.ai coding plan is it.

iFlow CLI: the aggregator that didn’t make it

iFlow CLI deserves a short obituary rather than a review. Through 2025 it was the wave’s tasting menu — free access to Qwen3-Coder, K2, DeepSeek, and GLM classes behind one terminal agent, the cheapest possible way to A/B the Chinese models. Per mid-2026 reporting, iFlow CLI has shut down.

The lesson generalizes: free multi-model aggregation had no visible business model, and in the year of the great harness die-off that was fatal. If iFlow was your sampler, the replacements are unglamorous — an OpenRouter-style aggregator with a real business model, or an open-source harness plus one API key per vendor.

Where DeepSeek fits: the model without a flagship CLI

The asymmetry is almost funny: the lab that started the wave still ships no flagship first-party CLI. A DeepSeek-TUI circulates in mid-2026 harness maps, but it reads as an ecosystem artifact rather than a strategic product. DeepSeek reaches coding agents the infrastructure way — through open-source harnesses and compatible endpoints.

The current line makes that posture stronger, not weaker. DeepSeek V4 went GA July 19, 2026; V4 Flash lists at $0.14 per million input and $0.28 per million output, the credible price floor for agentic coding, while V4 Pro carries the capability flag at a vendor-reported ~80.6% SWE-bench Verified. Meanwhile V3 and R1 — the models that made the lab famous — were deprecated on July 24, 2026, which quietly broke every config still pointing at the old IDs. DeepSeek’s docs describe an Anthropic-format endpoint usable straight from Claude Code, which made it the cheapest plausible engine for the plug-in pattern long before the flat plans arrived.

Strategy, in one line: DeepSeek monetizes tokens and mindshare, not seats — the research-lab posture we unpacked in the DeepSeek economics piece, unchanged even as everyone else built storefronts.

The plug-in pattern: cheap plans driving US harnesses

Here is the power move the whole article has been building to, and it survived every shutdown and repricing of 2026 intact. Picture a Friday-afternoon bulk refactor: forty files, mechanical changes, nothing subtle. You point Claude Code at the GLM Coding Plan’s endpoint and let the Lite tier eat the grind, while your Anthropic subscription stays fresh for the review pass and the two genuinely hard functions. Same harness, same muscle memory, different engine per task — that is the arbitrage in one image.

The mechanics are almost anticlimactic. Per the mid-2026 harness map, every serious CLI harness now accepts at least one of OpenAI-compatible or Anthropic-Messages endpoints, so the wiring is an environment-variable swap: a base URL and an auth token pointed at a vendor’s compatible endpoint. Z.ai, Moonshot AI, and DeepSeek all document the Claude Code swap officially. Router tools (claude-code-router, LiteLLM-class proxies) let you switch per project or per prompt, and open-source harnesses take provider config natively.

The plug-in pattern: Chinese coding plans driving US harnesses through compatible endpoints One env-var swap: cheap plans on the left, the harnesses you already use on the right.

Two sober notes before you wire anything. First, terms: harness licenses and provider ToS treat third-party backends differently, and “documented compatible” is not “endorsed” — read both sides before production use. Second, sprawl: every added provider is another account, another key, another quota window, another billing surface, and running that fleet without chaos is its own discipline.

Product note: The plug-in pattern multiplies accounts, keys, and endpoints fast. Automater Lite’s Toolbelt manages provider accounts (OAuth/API keys) across your CLIs, and local per-provider token metering shows what each lane actually burns — one ledger for a fleet that spans continents. Free on automater.ai.

The recipes: wiring GLM, Kimi, DeepSeek, and Qwen Code

Four recipes cover most of the pattern. Each snippet is the documented shape at publish; endpoints and model IDs churn, so verify against the vendor’s current guide before trusting a long run to it.

1. Claude Code → Z.ai GLM Coding Plan. The canonical swap, straight from Z.ai’s setup guide:

export ANTHROPIC_BASE_URL="https://api.z.ai/api/anthropic"
export ANTHROPIC_AUTH_TOKEN="<your-z.ai-key>"
export ANTHROPIC_MODEL="glm-5.2"
claude

Known failure mode: the five-hour quota window resetting mid-run — a long agent session that starts near the window’s edge stalls halfway through a refactor.

2. Claude Code → Kimi For Coding. Moonshot’s documented Anthropic-compatible endpoint:

export ANTHROPIC_BASE_URL="https://api.moonshot.ai/anthropic"
export ANTHROPIC_AUTH_TOKEN="<your-moonshot-key>"
export ANTHROPIC_MODEL="kimi-k3"
claude

Known failure mode: preview-era limits that move without notice; pin your expectations to the dashboard, not to a blog post — including this one.

3. Claude Code → DeepSeek V4. The cheapest lane:

export ANTHROPIC_BASE_URL="https://api.deepseek.com/anthropic"
export ANTHROPIC_AUTH_TOKEN="<your-deepseek-key>"
export ANTHROPIC_MODEL="deepseek-chat"
claude

Known failure mode: stale model IDs. Configs that still name V3-era models have returned errors since the July 24 deprecation; audit anything you wired before summer.

4. The open-source lane. OpenCode, Aider, Crush, and friends take any OpenAI-compatible provider in config — for example, Aider against Qwen3-Coder-Next on Alibaba Cloud:

export OPENAI_API_BASE="https://dashscope-intl.aliyuncs.com/compatible-mode/v1"
export OPENAI_API_KEY="<your-alibaba-cloud-key>"
aider --model openai/qwen3-coder-next

Known failure mode: per-family tool-call quirks — each model family formats tool calls slightly differently, and a harness tuned for one can mis-parse another until you set the right provider profile.

Capability honesty: where these models compete — and where they trail

The numbers, labeled by who produced them, as of August 2026. Kimi K3’s ~93.4% SWE-bench Verified is independent (Vals AI). GLM-5.2’s ~78.7% Verified is independent (Epoch AI); its 62.1% SWE-bench Pro figure is vendor-reported. DeepSeek V4 Pro’s ~80.6% Verified and Qwen3-Coder-Next’s ~70.6% are vendor-reported. On scoped agentic work — the kind benchmarks measure — the top of this wave now sits at or inside frontier range, which was not true eighteen months ago.

The same shape repeats beyond SWE-bench: on terminal-agent benches and tau-class tool-use evals, the current Chinese flagships sit within striking distance of frontier on scoped agentic work. Scoped is the load-bearing word.

Keep two discounts loaded. Vendor scores ride bespoke harnesses and retry budgets, so single-digit gaps are noise until reproduced like-for-like — and few vendors disclose their scaffolding in enough detail to reproduce anything. And benchmarks are scoped by construction: the place practitioner reports still consistently favor frontier models is long-horizon autonomy — multi-hour runs that must notice and recover from their own mistakes.

So route accordingly. The value is real for bulk implementation, test generation, refactors, and triage — high-volume work where the plan quota, not per-token price, is the budget. Frontier still earns its price on gnarly debugging, architecture calls, and anything unattended overnight. And no chart — ours included — substitutes for running your own evals on your own tasks before real work moves lanes.

The trust ledger: jurisdiction, terms, and the three paths

The trust question around Chinese AI coding tools usually arrives as one blunt question and deserves three separate answers, because there are three distinct paths, and conflating them produces bad decisions in both directions.

Three trust paths for Chinese AI coding tools: first-party API, Western-hosted weights, self-hosted Same weights, three very different contracts. Pick the path per workload, not per headline.

Path 1 — first-party APIs and plans. Your prompts and code transit vendor infrastructure under Chinese-jurisdiction or affiliated-entity terms. Before wiring one in, read four things: data residency, training-on-inputs defaults and opt-outs, retention windows, and governing law plus the actual contracting entity. This is the path the $3 plans live on; the savings are real and so is the homework.

Path 2 — Western-hosted open weights. Together-and-Fireworks-class hosts and OpenRouter-style aggregators serve the same open weights under the host’s terms and a US or EU jurisdiction. You give up flat-plan pricing, keep most of the price advantage, and move the contract to a counterparty your legal team has probably already reviewed.

Path 3 — self-hosted open weights. No transit at all; the model license is the only contract, and the MIT and Apache-2.0 licensing across this wave is about as clean as licensing gets. The hardware bar keeps dropping — the open-weight models guide covers what actually runs where.

Two compliance realities, stated without drama. Export-control and entity-list status varies by vendor and can change; employer procurement rules may decide the question regardless of your personal read. The honest advice is not “safe” or “unsafe” — it is: know your path, know your data classification, and know what your org allows.

Who should ride the wave — and how

The fit lines are clean. Strong fit: solo developers, side projects, and high-volume grind work where quota beats quality-ceiling; teams comfortable operating on paths 2 or 3. Poor fit: anyone under data-residency or procurement constraints who cannot self-host — for you the answer is path 3 or nothing, and that is a legitimate answer.

The operating pattern that works is the budget-grind lane: a cheap plan (GLM Lite-class, Kimi For Coding) or floor-priced API (DeepSeek V4 Flash) carries the mechanical bulk, your premium harness and model carry the judgment steps, and evals gate what graduates from one lane to the other.

The adoption sequence, updated for the post-free-tier era: start metered, not subscribed — a week of real work through DeepSeek V4 Flash or the Qwen API costs about a coffee and produces an actual volume number. Buy one flat plan only after that week proves the volume supports it. Wire it into your primary harness last, keep the wiring in version control as reviewable config, and expect to revisit it monthly — this market repriced, renamed, or retired something every month of 2026 so far.

Here is that first week, concretely. Day one: create one metered account, wire it into a spare harness profile, and re-run yesterday’s actual tasks side by side. Days two through four: route only bulk steps — test generation, mechanical refactors, triage summaries — through the cheap lane and keep judgment work where it lives today. Day five: tally the meter, count the interventions, and decide one of three ways — stay metered, buy a flat plan because volume justifies it, or walk away having spent less than lunch. Any of the three is a good outcome, because all three replace a vibe with a number.

FAQ: Chinese AI coding tools

What is Qwen Code?

Qwen Code is Alibaba’s open-source agentic coding CLI, forked from the now-discontinued Gemini CLI and retuned for the Qwen3-Coder model line, shipped under Apache-2.0. Its famous free OAuth tier — roughly 2,000 requests a day — ended April 15, 2026; today you bring an Alibaba Cloud or compatible API key.

Is the GLM Coding Plan really $3 a month?

At the promotional Lite rate, yes, as marketed. List pricing runs higher, quotas are metered in prompts per five-hour window, and the “multiples of Claude” framing is a vendor claim rather than a measurement. Read the renewal terms before assuming the promo price is the permanent price.

Can I use Chinese models inside Claude Code?

Mechanically, yes. Z.ai, Moonshot AI, and DeepSeek all document Anthropic-compatible endpoints, so a base-URL and auth-token swap points Claude Code at GLM-5.2, Kimi K3, or DeepSeek V4. Check both vendors’ terms first — compatibility being documented is not the same as the pairing being endorsed.

Is it safe to send code to Chinese AI APIs?

It depends on the path. First-party APIs and plans route code through vendor infrastructure under Chinese-jurisdiction terms; Western-hosted open weights move the contract to a US or EU host; self-hosting removes transit entirely. Classify your data, read the terms for your chosen path, and check employer policy.

Sources