Kimi K3, GLM-5.2, DeepSeek V4: The Open Models Crowding the Frontier
Kimi K3, GLM-5.2, DeepSeek V4, Qwen3-Coder-Next: verified figures, prices, deployment lanes, and how to pick the best open source model 2026 for agent work.
Go deeper. Build your own.
On July 16, 2026, an open-weight model took the top slot on a frontend-code arena ahead of Claude Fable 5 — five weeks after Fable 5 launched as Anthropic’s Mythos-class flagship. That had never happened before. Not “an open model got close.” An open model, Moonshot AI’s Kimi K3, ranked first, with the reigning frontier model underneath it.
If you are hunting the best open source model 2026 has produced for agent work, the honest answer changed four times between April and July: DeepSeek V4 on April 24, GLM-5.2 on June 16, Kimi K3 on July 16, and Qwen3-Coder-Next holding down the only lane the others can’t reach — your own hardware. Each release moved a different axis: benchmark ceiling, price floor, independent verification, local deployability.
This is the scorecard. For each of the four you get parameters, license, context, benchmark results with who actually ran them — vendor-run and independent numbers are not the same currency, and we label every figure — plus current prices and where the weights really are. Then the parts a spec sheet won’t tell you: what the K3-beats-Fable-5 headline does and doesn’t mean, the three deployment lanes and their trust trade-offs, how to eval before you move traffic, and where this goes next.
Ninety days that rearranged the leaderboard
The compressed timeline is the story. Anthropic shipped Claude Fable 5 on June 9. Within six weeks, three Chinese labs shipped open-weight responses that bracketed it from below — and one of them touched it.
| Date | Event |
|---|---|
| April 24, 2026 | DeepSeek V4 first release (Pro and Flash lines) |
| June 9, 2026 | Claude Fable 5 launches — the frontier reference point |
| June 16, 2026 | GLM-5.2 ships under MIT; #1 open model on Artificial Analysis at launch |
| July 16, 2026 | Kimi K3 ships; tops a frontend-code arena ahead of Fable 5 |
| July 19, 2026 | DeepSeek V4 general availability |
| July 24, 2026 | DeepSeek deprecates V3 and R1 on the first-party API |
Two patterns make this cohort different from earlier open-weight waves. First, convergence: all four are sparse mixture-of-experts models, all four ship 1M-token context windows or close to it, and all four were post-trained hard on agentic tool use rather than chat. These are models built to sit inside a harness loop, not a chatbox. Second, verification grew up. Independent evaluators — Vals AI, Epoch AI, Artificial Analysis — now re-run the headline benchmarks within days of release, which means we can separate what a vendor claims from what a referee reproduced.
That separation is this scorecard’s method. Vendor-run means the lab scored its own model, with its own harness, and published the number. Independent means a third party reproduced the run. Both are useful; only one is evidence you should re-route production on. Where a figure below is vendor-run, it says so.
The best open source models of 2026, at a glance
| Kimi K3 | GLM-5.2 | DeepSeek V4 Pro / Flash | Qwen3-Coder-Next | |
|---|---|---|---|---|
| Lab | Moonshot AI | Zhipu (Z.ai) | DeepSeek | Alibaba (Qwen) |
| Released | Jul 16, 2026 | Jun 16, 2026 | Apr 24, GA Jul 19, 2026 | 2026 line, current |
| Parameters | 2.8T MoE | 744B / 40B active | 1.6T / 49B · 284B / 13B | 80B / ~3B active |
| Context | 1M | 1M | 1M | long-context |
| SWE-bench Verified | ~93.4% — independent (Vals AI) | ~78.7% — independent (Epoch AI) | ~80.6% (Pro) — vendor-run | ~70.6% — vendor-run |
| License | modified MIT expected | MIT | MIT | Apache-2.0 |
| Weights | staged; not yet posted | posted | posted | posted |
| Price (per M in/out) | first-party API | ~¼ of frontier output price | Flash: $0.14 / $0.28 | $0.11 / $0.80 (Alibaba Cloud) |
| The claim it owns | benchmark ceiling | verified value | price floor | runs on one box |
Figures compiled from the mid-2026 open-model roundup at Morph, the labs’ own release materials, and the independent runs noted. As of August 2026; every number here has a shelf life measured in weeks.
The four cards side by side. Green scores were independently reproduced; amber scores are the vendor grading its own homework.
Kimi K3: the ceiling, on a promise
Moonshot AI shipped Kimi K3 on July 16 as a 2.8-trillion-parameter MoE with a 1M-token context window, and it is the open model with the loudest claim: ~93.4% on SWE-bench Verified — and that number is independent, from Vals AI’s re-run, not Moonshot’s deck. Add the frontend-code arena result, where K3 became the first open model to lead one ahead of Claude Fable 5, and you have the strongest benchmark case any open release has ever made.
The post-training story explains the shape of the results. Moonshot’s release materials describe K3 as tuned on agentic trajectories — multi-step tool use, edit-run-fix loops — rather than single-shot completion , which would explain strength in exactly the places harness users notice: patch discipline, tool-call formatting, recovering from a failed test run instead of thrashing.
Now the asterisk, stated plainly: as of late August 2026, the weights are staged, not posted. Moonshot launched the API first, with weights promised to follow under an expected modified-MIT license. Until they land, K3 is functionally a closed model with an open roadmap — you can buy it as a service from Moonshot, but you cannot re-host it, audit it, or run it under terms you control. That matters for the deployment lanes below, and it is why our scorecard lists “the claim it owns” as the ceiling rather than openness.
Where to run it today: Moonshot’s first-party API, which exposes an Anthropic-compatible endpoint that drops into Kimi Code CLI and other Anthropic-native harnesses. Expect Western re-hosts within days of the weights actually posting — that is the pattern GLM and DeepSeek established.
GLM-5.2: the value line with independent receipts
Zhipu’s Z.ai shipped GLM-5.2 on June 16: 744B total parameters, 40B active, MIT license — the real MIT, no scale carve-outs — and 1M context. At launch it took the #1 open-model slot on Artificial Analysis’s intelligence index, and its benchmark story is the most instructive of the summer because it comes in two numbers that look contradictory and aren’t.
Zhipu’s own headline is ~62.1% on SWE-bench Pro, the harder successor benchmark built after the original started saturating. Epoch AI’s independent reproduction on classic SWE-bench Verified landed at ~78.7%. Different benchmarks, different difficulty, both real: the vendor chose to advertise the harder test, and a referee confirmed strong results on the standard one. When a lab’s independently-verified number is the flattering one, that is the trust signal — the reverse pattern is the one to squint at.
The commercial context does real work here too. GLM-5.2 is the engine behind the Z.ai GLM Coding Plan, priced at roughly a quarter of frontier output rates on the API. Zhipu’s subscription revenue depends on this model completing tool calls inside real harnesses every day — a direct commercial incentive to keep function calling reliable at this price tier. MIT weights mean every deployment lane is open: first-party API, Western re-hosts, or your own cluster — though at 744B total, that last one is a lab project, not a laptop one.
DeepSeek V4: the price floor, in two sizes
DeepSeek released V4 on April 24 and took it to general availability on July 19 — then, five days later, deprecated V3 and R1 on the first-party API. If your configs still name V3-era models, that is not a scorecard question, it is an outage; our DeepSeek V4 migration guide covers the swap, the eval sequence, and the rollback plan.
The V4 line splits deliberately. V4 Pro is the capability play: 1.6T total parameters, 49B active, ~80.6% on SWE-bench Verified — vendor-run, so calibrate accordingly; no independent referee had published a full reproduction as of late August. V4 Flash is the economics play: 284B total, 13B active, priced at $0.14 per million input tokens and $0.28 per million output — the credible price floor for agentic coding, full stop. Both are MIT-licensed with 1M context, and both post weights for real, which is why third-party hosts already serve them.
Flash is the number that bends behavior. Agent loops re-send growing context every step, so an agent task burns tokens at 10–100x chat rates — at $0.14, workload categories that were not viable at frontier prices become rounding errors. We work that arithmetic in our analysis of open models and collapsing agent costs; the short version is that Flash made “just let the agent grind on it” a defensible default for mechanical work.
Qwen3-Coder-Next: the frontier you can carry
Alibaba’s Qwen3-Coder-Next owns the lane none of the giants above can enter: it runs on hardware you own. The full model is an 80B-parameter MoE with only ~3B active per token, Apache-2.0 licensed, and it fits in roughly 46GB of unified memory — a 64GB Mac, a Ryzen AI Max box, a DGX Spark. The 30B Flash variant squeezes into ~18GB, which puts it on a single 24GB GPU. If the phrase “frontier-adjacent model on one machine” appears in a 2026 argument, this model is the evidence; we spec the boxes it runs on in the local AI workstation guide.
The benchmark figure is ~70.6% SWE-bench Verified — vendor-run, no independent reproduction yet — which would have led the entire open field a year ago and now ranks fourth of four in this cohort. That is the correct context in both directions: the field moved fast, and 70% Verified from 3B active parameters running beside its KV cache in local silicon is still a little absurd.
Know its shape before you rely on it. Three billion active parameters make Qwen3-Coder-Next a superb executor and a mediocre architect: fast, cheap, disciplined on scoped tasks, weaker on multi-hour planning. Hosted via Alibaba Cloud it costs $0.11/$0.80 per million tokens; the QwenLM GitHub organization and the Qwen collection on Hugging Face are where quantizations and serving fixes land first — one of the most-forked open bases means fast community patches when something breaks.
What “K3 beat Fable 5” does — and doesn’t — mean
The headline deserves surgical handling, because both the hype and the dismissal get it wrong.
What it does mean. An open-weight lab shipped a model that leads a public arena over the flagship Anthropic released five weeks earlier. The open-to-frontier gap — commonly put at a year-plus in 2023 and months in 2024 — is now measured in weeks on some axes, and open models are benchmarked against the frontier flagship rather than against each other. That reframes what “frontier” means: increasingly it names a price band and a trust posture, not a capability monopoly.
What it doesn’t mean. One frontend arena is one distribution: preference-voted, short-horizon, heavy on visible UI output — the genre open models have always punched hardest in. It is not a measure of eight-hour autonomous refactors, recovery from ambiguous specs, or the long-tail judgment calls that separate flagships in daily driving. SWE-bench Verified itself is saturating — when the top of the field posts 78–93%, the benchmark stops discriminating precisely where you need it to — and harness effects can move any model’s score by double digits. We unpack how to read all of this in our guide to agent benchmarks in 2026.
Our read, labeled as such: for long-horizon agent work, Fable 5 remains ahead on the axes arena votes don’t capture, and the gap is real but no longer comfortable. The correct response to the headline is neither to switch nor to scoff — it is to eval, which is a section below.
Three lanes to run them — and what each asks you to trust
Every open-weight model reaches production through one of three lanes, and the lane decides your price, your latency to new releases, and what you are trusting.
Lane 1: first-party APIs. Moonshot, Z.ai, DeepSeek, Alibaba Cloud. Cheapest per token, first to serve each release, and the lane with the most to think about: your prompts and outputs transit the lab’s infrastructure under its jurisdiction’s terms. Fine for open-source work and public repos; a real conversation with legal for anything else. The full trust checklist — data terms, retention, procurement — is in our review of Chinese frontier models for agent work.
Lane 2: Western re-hosts. Together, Fireworks, OpenRouter, and the hyperscaler catalogs serve MIT and Apache weights under their own terms of service. You pay a multiple of first-party pricing — still far under closed-frontier rates — and in exchange you get contracts you can sign, data residency you can name, and SLAs. This lane only exists for models whose weights actually posted, which today means GLM-5.2, V4, and Qwen3-Coder-Next, and not yet K3.
Lane 3: self-host. Full control, no per-token meter, and a hardware bill. Qwen3-Coder-Next is the realistic one-box option; GLM-5.2 and V4 are multi-GPU cluster projects; K3 is nobody’s self-host until the weights exist. Serving stacks, quantization trade-offs, and harness wiring live in our guide to open-weight models that can drive a harness.
The lane you pick sets price, release latency, and who you trust. Only models with posted weights can leave lane one.
Eval before you switch: the two-week bake-off
A scorecard tells you which models deserve a trial. It cannot tell you which one survives contact with your repos, your harness, and your definition of done. Before any of these four takes production traffic, run the bake-off:
- Assemble 20–50 golden tasks from your own recent work — real tickets, real repos, tasks where you know what a good outcome looks like. Public benchmarks are the models’ home turf; your backlog is yours.
- Run them through your actual harness, not a raw API playground. Harness effects are large enough to reorder this scorecard, and the model you deploy will live inside whichever tool from the agent harness field map you drive daily.
- Score four things per task: solve rate, cost per solved task (not per token — a cheap model that needs three attempts isn’t cheap), wall-clock time, and diff quality on human review.
- Compare against your incumbent’s baseline on the same tasks, and hold the trial for two weeks of real ad-hoc use before moving traffic. This is the same daily-driver standard we apply in our best agentic AI tools testing — a model that aces the golden set but annoys you for ten working days has told you its answer.
Product note: The hardest part of the bake-off is step one — most teams don’t have a task set that mirrors their real work, because their real work evaporated with each closed terminal. Automater Lite archives every session from 10+ providers (Claude Code, Codex, Qwen Code, Kimi Code CLI, OpenCode, and more) into one searchable, local-first library. When a new model ships, last month’s sessions are a ready-made eval set: search, export, replay. Free on automater.ai.
Where this goes next
Predictions, dated August 2026, so you can grade us later.
Staged weights become the default release pattern. K3 formalized what the summer hinted at: API launch first, weights after the news cycle peaks. Expect the gap between “released” and “downloadable” to become a standard two-to-eight-week window — and treat any scorecard that ignores the distinction as marketing.
The floor drops again. $0.14 input is not the bottom. Between V4 Flash’s trajectory, Qwen’s $0.11 hosted rate, and the ~80%-per-year open-model price declines reported through mid-2026, a sub-$0.10 agentic-grade input price within two quarters is the safe bet. Budget models, not budgets, are the constraint now.
A US open-weight answer, sized mid, not max. OpenAI’s gpt-oss line proved US labs will ship open weights when the strategic math demands it; a follow-up aimed at the Qwen3-Coder-Next one-box class — where adoption is stickiest — is the likelier response than an open 2.8T flagship.
Re-run this scorecard quarterly. Every number here was set between April and July of one summer. The models are converging on the same recipe — sparse MoE, 1M context, agentic post-training, permissive license — which means leapfrogging is cheap and rankings are perishable. Your routing table should have a review date on it.
FAQ: the best open source models of 2026
What is the best open source model for coding in 2026?
There is no single answer — pick by constraint. Kimi K3 holds the benchmark ceiling (~93.4% SWE-bench Verified, independently run by Vals AI). DeepSeek V4 Flash holds the price floor at $0.14/$0.28 per million tokens. Qwen3-Coder-Next is the strongest model that runs on one machine. GLM-5.2 balances verified capability and quarter-of-frontier pricing.
Did Kimi K3 really beat Claude Fable 5?
On one frontend-code arena, yes — K3 became the first open-weight model to lead an arena ahead of Fable 5 in July 2026. That is a genuine milestone, not an overall verdict: arena votes measure short, visible tasks, and Fable 5 still leads on long-horizon agent reliability in most practitioners’ testing.
What is the cheapest model for agentic coding in 2026?
DeepSeek V4 Flash, at $0.14 per million input tokens and $0.28 per million output on the first-party API, is the credible price floor as of August 2026. Alibaba’s Qwen3-Coder-Next is nearby at $0.11 input (though $0.80 output) via Alibaba Cloud, and it can also run locally for zero marginal cost.
Which open-weight model can I run on my own machine?
Qwen3-Coder-Next is the realistic one-box option: the full 80B MoE fits in about 46GB of unified memory, and the 30B Flash variant runs in roughly 18GB — one 24GB GPU. GLM-5.2 (744B) and DeepSeek V4 (284B–1.6T) need multi-GPU clusters; Kimi K3 has no posted weights yet.
Are Kimi K3’s weights actually released?
Not as of late August 2026. Moonshot AI launched K3’s API on July 16 with weights staged to follow under an expected modified-MIT license. Until they post, K3 can only be used through Moonshot’s API — no re-hosting, no self-hosting, no independent audits of the weights themselves.
Sources
- Best open-source coding models of 2026 — Morph
- Vals AI — independent model evaluations
- Epoch AI — independent benchmark reproductions
- Artificial Analysis — model intelligence index
- Moonshot AI — Kimi models and platform
- Z.ai — GLM models and plans
- DeepSeek — models and API
- Qwen on Hugging Face — open-weight checkpoints
- QwenLM on GitHub — Qwen3-Coder repositories
- Introducing Claude Fable 5 and Claude Mythos 5 — Anthropic
