Chinese AI Models for Agent Work: DeepSeek, Qwen, Kimi, GLM
Kimi K3, GLM-5.2, DeepSeek V4, and Qwen3-Coder-Next: a lab-by-lab guide to Chinese AI models for agent work, covering capability, licenses, access, and trust.
Go deeper. Build your own.
Pull up any open-model leaderboard in August 2026 and the top of it reads like a Shanghai tech-park directory: Kimi K3 at roughly 93% on SWE-bench Verified, GLM-5.2 leading the Artificial Analysis open index, DeepSeek V4 selling tokens at fourteen cents a million. Chinese AI models no longer just set the price floor for agent work — among open-weight models, they increasingly set the capability ceiling too.
This is a working survey for people who route real agent workloads and need to decide what to route them to. The economics history lives in our DeepSeek economics analysis; the geopolitics lives elsewhere. Here, every lab gets the same single question: can its model do agent work, and on what terms — capability, license, access path, data terms.
The cast: Moonshot AI’s Kimi K3, Zhipu’s GLM-5.2, DeepSeek’s V4 pair, Alibaba’s Qwen3-Coder-Next, then MiniMax and a second tier worth a line each. Facts verified as of August 27, 2026; this field reshuffles quarterly.
The labs setting the floor
Two years ago the honest framing was “impressive for the price.” That framing is dead. The current generation competes on capability first — a Chinese open-weight model now leads at least one coding arena outright ahead of Claude Fable 5 — while the price gap somehow widened, with open-model API pricing down roughly 80% year over year per the mid-2026 harness map.
What did not change is the strategy: ship open weights under permissive licenses, sell cheap first-party access, and let the ecosystem — hosts, harnesses, fine-tuners — do the distribution. Every profile below is a variation on that playbook, and the variations are where the routing decisions hide.
What agent work demands of Chinese AI models
Agent fitness is not chat quality. A model that writes a charming explanation can still be a bad agent, because harness work stresses different muscles: tool calls that stay well-formed across dozens of sequential invocations, instructions that survive a long horizon without drift, context handling under the constant re-sending an agent loop produces, and SWE-bench-class coding ability as the baseline entry fee.
Coding benches are necessary but not sufficient here. SWE-bench-class results measure scoped problem-solving; tau-class tool-use evals and terminal-agent benches stress the loop itself, and long-horizon behavior — hour three of an unattended run — is still measured mostly by practitioner scar tissue. All three layers appear in the profiles below, labeled.
Then come the terms, which fail more adoptions than capability does. Three dimensions matter: the weights license (what you may ship), the access paths (first-party API, Western host, self-host), and the data terms attached to each path. A model with frontier scores and unacceptable terms for your workload fails the test; a slightly weaker model you can self-host under MIT may pass it.
One discipline runs through everything below: lab-reported numbers and independent evaluations are labeled as such, everywhere. And no label substitutes for running your own evals on your own tasks — the governing rule for every adoption decision in this survey.
The road here: K2, GLM-4.x, and the V3 era
The current lineup is one summer old, so a compressed history earns its space. DeepSeek V3 arrived in December 2024 with near-frontier capability and a training bill reported in single-digit millions; R1 followed in January 2025 and made “reasoning model, MIT license, open weights” a sentence that moved markets. Through 2025 DeepSeek kept shipping sparse-attention efficiency work that cut serving costs again and again.
July 2025 brought Kimi K2 — a trillion-parameter MoE with 32B active, post-trained explicitly on synthesized tool-use data, the first open-weight model whose pitch was “agent” rather than “chat.” Zhipu’s GLM-4.5 and 4.6 built the value-coding franchise the same year: MIT weights, Anthropic-compatible endpoints, and coding plans priced like a streaming subscription. Alibaba’s Qwen3 family, Apache-2.0 from edge sizes to flagship MoE, became the most-forked base in open modeling — DeepSeek’s own R1 distills rode on Qwen bases.
That era closed fast. In 2026 each lab superseded its 2025 headliner, and DeepSeek formally deprecated V3 and R1 on July 24, 2026 — a reminder that “open weights” does not mean “your API config is immortal.” The four profiles that follow are the models that replaced them, per the 2026 open coding-model roundup. Anyone still quoting a K2 or GLM-4.6 benchmark is describing the previous generation; searches for kimi k2 now mostly deserve a K3 answer.
Kimi K3: the agent-first flagship
Kimi K3 is Moonshot AI’s flagship MoE, released July 16, 2026: 2.8T total parameters, a 1M-token context window, and the strongest independent agentic-coding result of any open-weight model — roughly 93.4% on SWE-bench Verified per Vals AI. It extends K2’s defining bet, post-training aimed squarely at tool-use reliability over long call sequences.
The line. K3 headlines; K2 remains available and widely hosted. The K-line’s signature is agentic post-training at scale — tool-call discipline as a first-class training objective, with thinking variants for long multi-step runs.
Agentic strengths. The SWE-bench Verified number is independent, not vendor marketing, and K3 became the first open-weight model to lead a frontend-coding arena ahead of Claude Fable 5. Practitioner reports emphasize call-format stability deep into long sessions — precisely the property harness work punishes models for lacking.
Terms. A modified-MIT license is expected, following K2’s pattern: permissive, plus an attribution condition that binds only at very large commercial scale. Weights were staged to land after the API launch, so check release status before planning self-hosting. Access today runs through Moonshot AI’s API and the Kimi For Coding plan, with Western hosts expected to onboard once the weights actually post.
The honest weakness. Availability lag. Until the weights fully land, K3 is effectively a first-party-API model wearing an open-weight badge — path-2 and path-3 users are still waiting, and a 2.8T MoE will be a heavyweight host even then.
GLM-5.2: the value line grows teeth
GLM-5.2 is Zhipu’s current flagship, published under the Z.ai brand on June 16, 2026: 744B total parameters with 40B active, MIT-licensed, 1M context, and the #1 open-weight model on the Artificial Analysis intelligence index. The GLM model line used to be the value pick; 5.2 makes it a capability pick that happens to stay cheap.
The line. GLM-5.2 succeeds the 4.5/4.6 coding era. First-party output pricing runs around a quarter of frontier rates, and the Z.ai coding plans repackage it at flat monthly prices.
Agentic strengths. Independent results back the promotion: ~78.7% on SWE-bench Verified per Epoch AI, alongside a vendor-reported 62.1% on the harder SWE-bench Pro. It is built for harness duty — Anthropic-compatible endpoints, official Claude Code setup guides, and the whole plan apparatus covered in our Chinese CLI wave review.
Terms. MIT weights — the cleanest license in the class — plus first-party API, coding-plan endpoints, and broad Western-host availability. Zhipu’s Tsinghua lineage and its public-listing status matter here only as stability signals for the plan you might subscribe to.
The honest weakness. The Pro-vs-Verified gap. A 16-point drop from Verified to the vendor’s own SWE-bench Pro number says the ceiling is real on the hardest tickets; K3 and frontier models hold up better as tasks get genuinely nasty.
DeepSeek V4: Pro, Flash, and the price floor
DeepSeek V4 is the efficiency pioneer’s 2026 line, released April 24 and GA July 19: V4 Pro at 1.6T total parameters with 49B active, and V4 Flash at 284B with 13B active. Both are MIT-licensed with 1M context. The DeepSeek models remain the reference answer to “how cheap can real agent work get.”
The line. Pro carries capability — a vendor-reported ~80.6% SWE-bench Verified. Flash carries the market: $0.14 per million input tokens, $0.28 per million output, the credible price floor for agentic coding in 2026. The V3 and R1 era ended by decree on July 24, 2026, when DeepSeek deprecated both; if your configs predate summer, the V4 migration guide is the fastest path through the breakage.
Agentic strengths. Frontier-adjacent coding and reasoning at prices that make retry budgets and multi-sample strategies economically boring. Flash in particular changes agent design: at floor pricing, “run it three times and vote” is a rounding error.
Terms. MIT weights, the cheapest first-party API in the class, and the broadest Western-host coverage of any lab here. Self-hosting is realistic for Flash-class deployments and a multi-GPU project for Pro.
The honest weakness. Tool-calling reliability historically trailed the purpose-trained K-line, and independent evidence on whether V4 closed that gap is still accumulating — worth a targeted eval before you hand it a long unattended run.
Qwen3-Coder-Next: the one-box frontier
Qwen3-Coder-Next is Alibaba’s agentic-coding specialist and the survey’s outlier bet: 80B total parameters with only ~3B active, Apache-2.0, and a vendor-reported ~70.6% SWE-bench Verified — deliberately trading peak capability for the ability to run on one machine. Around 46GB of unified memory serves it; a 30B Flash variant fits in ~18GB.
The line. The Qwen models’ differentiator has always been breadth — edge sizes to flagship MoE under one Apache-2.0 umbrella — and Coder-Next is the breadth argument sharpened to a point: the frontier-adjacent model you can own outright on a workstation.
Agentic strengths. The active-parameter budget makes it fast and cheap everywhere it runs: $0.11 per million input and $0.80 per million output on Alibaba Cloud, or local inference with no meter at all. It is the most-forked ecosystem in open modeling, so quantizations, fine-tunes, and community fixes land here first — weights on Hugging Face and ModelScope, plus every major Western host.
Terms. Apache-2.0 with patent grant across the open line; closed API-only tiers (the Max class) are a separate product family.
The honest weakness. The scoreboard. A ~70.6% vendor-reported Verified is the lowest headline number of the four, and long-horizon consistency trails the trillion-class models. The trade is explicit: you are buying ownership and latency, not peak capability.
MiniMax and the second tier
The MiniMax model line leads the second tier: M-series MoEs built around compact activation — M2 shipped at 230B total with 10B active as last verified — with permissive licensing, aggressive pricing, and strong agentic scores at launch. It is the lab most likely to force its way into the top table by the next revision of this survey.
The rest of the tier, one line each, all worth a verify sweep at read time: ByteDance’s Seed/Doubao line is capable and mostly closed; Baidu opened its ERNIE 4.5 line under Apache-2.0; Tencent’s Hunyuan ships open weights across sizes.
This tier reshuffles quarterly, so treat the four profiles above as the stable spine and this section as the watchlist. Promotion off the watchlist takes the same yardstick as everything else: independent agentic-eval results, a permissive license, and Western-host pickup.
The agent-fitness comparison
The table below is the survey compressed. Vendor-reported numbers are marked (v); independent results are marked with their source. Verified as of August 27, 2026.
| Model (lab) | Released | Params (total/active) | SWE-bench Verified | Context | License | Cheapest serious path |
|---|---|---|---|---|---|---|
| Kimi K3 (Moonshot AI) | Jul 16, 2026 | 2.8T / n.p. | ~93.4% (Vals AI) | 1M | Modified MIT expected; weights staged | Kimi For Coding plan |
| GLM-5.2 (Zhipu / Z.ai) | Jun 16, 2026 | 744B / 40B | ~78.7% (Epoch AI); 62.1% SWE-bench Pro (v) | 1M | MIT | GLM Coding Plan Lite |
| DeepSeek V4 Pro | Apr 24 / GA Jul 19, 2026 | 1.6T / 49B | ~80.6% (v) | 1M | MIT | First-party API |
| DeepSeek V4 Flash | GA Jul 19, 2026 | 284B / 13B | n.p. — the speed/price tier | 1M | MIT | $0.14/$0.28 per M — the floor |
| Qwen3-Coder-Next (Alibaba) | 2026 | 80B / ~3B | ~70.6% (v) | n.p. | Apache-2.0 | Self-host at ~46GB |
Four model cards, one glance: capability, license, and the cheapest serious path to each.
Read it with the discount rule attached: vendor scores ride bespoke scaffolds and retry budgets, so single-digit gaps between labs are noise until reproduced like-for-like. The full standings, including the US open-weight entries, live in the mid-2026 open-weight scorecard.
The column the table cannot hold is long-horizon behavior, and the practitioner-report consensus runs like this. K3 holds tool-call formatting deepest into long sessions — that is what its post-training bought. GLM-5.2 stays solid on scoped runs and degrades first on the hardest multi-step tickets, consistent with its Pro-tier gap. V4 makes failure cheap rather than rare: retries cost so little that its recovery story is economic, not behavioral. Qwen3-Coder-Next drifts earliest in marathon sessions — the known price of a ~3B active budget. None of this transfers automatically to your tasks; a few benchmark points transfer even less.
The licensing map: what you can ship
| License class | Models | What it permits |
|---|---|---|
| MIT | DeepSeek V4 Pro and Flash; GLM-5.2 | Commercial use, modification, redistribution, no capability conditions — the maximal case |
| Apache-2.0 | Qwen3 open line incl. Qwen3-Coder-Next | Same freedoms plus an explicit patent grant |
| Modified MIT | Kimi K2; K3 expected | Fully permissive, plus an attribution condition that binds only at very large commercial scale |
Three license classes, and the rule that outlives them all: the weights license and the API terms are different contracts.
Calibrate against the industry’s looser usage of “open”: Llama-style community licenses carry acceptable-use policies and scale thresholds that permissive licenses simply don’t have. The Chinese labs’ licensing is, bluntly, cleaner — and that is a competitive choice, not an accident, because permissive weights are what let agent-product builders fine-tune, self-host, embed in paid products, and redistribute. The open-weight field guide maps those freedoms across every contender.
One rule matters more than any row of that table: the weights license and the API terms of service are different contracts. MIT weights say nothing about what a first-party API does with your prompts. Read both documents; they were written by different lawyers for different reasons.
Trust, by usage path: a checklist, not a verdict
Trust questions attach to whoever runs inference and under what terms — not to the weights’ country of origin. Open weights are static, inspectable artifacts; the meaningful variable is the path you use them through, so score these three paths per workload instead of debating the flag.
Path 1 — first-party APIs and plans. Chinese jurisdiction, each lab’s own retention and training-use terms. Cheapest access, heaviest homework.
Path 2 — Western-hosted open weights. Together-and-Fireworks-class hosts serve the same checkpoints under US or EU terms your legal team has likely already reviewed. You trade a little price for a lot of contractual familiarity.
Path 3 — self-hosted. Full control, your infrastructure risk, the license as the only contract; the open-source agent stack covers the serving layer that makes this real.
The checklist to run per workload: what data classification flows through the loop; what the retention and training-use terms say on your chosen path; which jurisdiction and compliance obligations apply; what audit and logging you need. Score each item per path, not per model — a public-repo refactor and a proprietary-core migration produce different scores on the same path, and export-control or employer procurement rules can settle the question before capability gets a vote.
The same model can be disqualified on path 1 and perfectly fine on path 3. The checklist makes that call — not this article, and not a headline.
Content behavior in agent workloads
Some checkpoints carry baked-in filtering on politically sensitive topics, documented by community testing. State it plainly and size it honestly: refactors, test generation, and tool calls essentially never touch that surface, which is why coding-agent users can go months without noticing it exists.
The exceptions cluster in content pipelines — news summarization, scraping steps, moderation work — where topic-filtered outputs have surprised builders: a summarization agent that goes quiet on a subset of articles fails silently, which is worse than failing loudly. Two mitigations: know your workload’s topic surface, and eval on real tasks. Note that Western hosting fixes data terms, not weight-level behavior — path choice and filtering are separate questions.
What the US response confirms
Watch what the incumbents do, not what anyone says. Since the R1 shock, US labs have cut prices, shipped cheap tiers, and deepened caching discounts ; OpenAI now publishes its own Apache-2.0 open-weight line in gpt-oss; and open-model API pricing across the market fell roughly 80% year over year. Competitors now price and publish against the floor these labs set — that is the market conceding the premise.
The correlation is stated honestly, the way our DeepSeek economics piece argues it: timelines side by side, causality probable but not proven. And the read is falsifiable — if US labs stop open releases, if the price gap closes, or if Chinese labs retreat from permissive licenses, this section is wrong and we will rewrite it.
Adopting Chinese AI models deliberately: the on-ramp
Step one — build the eval suite from your own work (half a day). Pull twenty real tasks from archived sessions, not synthetic benchmarks, and gate any switch on measured parity for the specific step being routed — a workable bar is matching the incumbent’s pass rate on that step within a margin you set in advance, not “felt fine.” The eval discipline is the whole game; nothing below matters without it.
Step two — red-team the lane (half a day). Run your prompt-injection and tool-abuse suite against the new model before it touches repositories. A new model lane is a new dependency and deserves a dependency’s scrutiny.
Step three — route bulk steps first (a metered week). Summarization, test grinding, boilerplate — high-volume, low-stakes, measurable. This is where floor pricing pays immediately, and where the best-tools rotation shows most daily drivers already quietly running a second lane.
Step four — keep the exit ramp (an hour, quarterly). OpenAI- and Anthropic-compatible endpoints make reversal a config edit. Re-benchmark quarterly, because every fact in this survey churns.
Product note: Your archived sessions are the eval set. Automater Lite archives sessions across 10+ CLIs and meters tokens per provider locally — replay real work against a new model lane and measure exactly what the switch saves. Free on automater.ai.
FAQ: Chinese AI models
What are the main Chinese AI models?
As of August 2026: Moonshot AI’s Kimi K3, Zhipu/Z.ai’s GLM-5.2, DeepSeek V4 in Pro and Flash tiers, and Alibaba’s Qwen3 family with Qwen3-Coder-Next as the agentic specialist. MiniMax’s M-series leads the second tier. All ship open weights under permissive or near-permissive licenses.
Is Kimi K2 open source?
Kimi K2 shipped under a modified MIT license: fully permissive, plus an attribution condition that binds only at very large commercial scale. Kimi K3 is expected to follow the same pattern, with weights staged after its API launch. Check the LICENSE file on the specific checkpoint you deploy.
Can I use Qwen commercially?
Yes. The open Qwen3 line, including Qwen3-Coder-Next, ships under Apache-2.0 — commercial use, modification, and redistribution are permitted, with a patent grant included. Alibaba’s closed API-only tiers are a separate product family. As always, read the specific checkpoint’s LICENSE file before shipping.
Which Chinese model is best for coding agents?
There is no single answer. Kimi K3 leads on independent SWE-bench Verified results (~93.4% per Vals AI), GLM-5.2 tops the Artificial Analysis open index, and DeepSeek V4 Flash wins on price. Route a week of your own tasks through each candidate before deciding anything.
Is it safe to use Chinese AI models for work?
Judge the path, not the flag. First-party APIs operate under Chinese-jurisdiction terms; Western hosts serve the same open weights under US or EU terms; self-hosting leaves only the license. Classify the workload, read the terms for your path, and follow your employer’s policy.
Do Chinese AI models censor content?
Some checkpoints carry baked-in filtering on politically sensitive topics, documented by community testing. Coding and tool-call workloads rarely touch that surface; content pipelines and news summarization occasionally do. Hosting location changes data terms, not weight behavior — so eval the filtering question on your real tasks.
Sources
- Best open-source coding models of 2026 — Morph
- Coding CLIs in mid-2026: the engineer’s map — dev.to
- Vals AI — independent model evaluations
- Epoch AI — independent benchmarking
- Artificial Analysis — model intelligence index
- Qwen on Hugging Face — open-weight checkpoints
- DeepSeek — models and API
- Z.ai — GLM models and plans
- Moonshot AI — Kimi models and platform
- Introducing gpt-oss — OpenAI
