Amp, Crush, and OpenClaw: The New-Wave Harnesses Worth Watching
Eight new AI coding tools mapped: Amp, Crush, OpenClaw, Ecodex, and more — with evidence tiers, survivor criteria, and a safe two-week trial protocol for 2026.
Go deeper. Build your own.
Across the spring and summer of 2026, at least five coding harnesses died — Gemini CLI loudest among them, taking CI pipelines down with it on June 18 — and the shelf refilled before the funerals ended. The most-cited census of what filled it, the mid-2026 harness map, lists a cohort of newcomers that keeps surfacing in threads: Amp, Crush, OpenClaw, Kilo CLI, Hermes Agent, DeepSeek-TUI, Ecodex, and Pi.
This is the field report on those eight — the new AI coding tools of 2026 that might deserve a slot on your machine, profiled with something most roundups skip: evidence grading. Some of these projects have public repos and vendor sites you can inspect. Others exist, for now, mostly as entries on one well-sourced map. We say which is which every single time, because the year of the die-off is exactly the wrong year to install a harness on vibes.
You’ll get mini-profiles with tiers and verification flags, an honest sort of what’s genuinely new versus rebadged, each newcomer scored against the survivor criteria the die-off taught us, a who-should-try-what, and the safe-trial protocol we’d use before letting any of them near a real repo.
Where the new AI coding tools of 2026 came from
The die-off cleared shelf space, and it also set the entry bar. This cohort was born after Claude Code defined the category, so the table stakes came pre-installed: MCP support, permission gates, and multi-model plumbing on day one. Per the mid-2026 map, every serious CLI harness now accepts at least one of the OpenAI-compatible or Anthropic-Messages endpoint formats , which means no newcomer wins by merely connecting to models. They have to differentiate on an organ.
And they do — that’s the cohort’s signature, visible across the 2026 harness field map: each newcomer bets on a single subsystem. Amp on context assembly, Crush on the terminal surface, Ecodex on the permission model, DeepSeek-TUI on unit economics. Nobody is trying to out-generalize the incumbents; the daily-driver rosters already have generalists.
A method note before the profiles, because this piece leans on it. Claims below carry one of two tiers. Documented means a public repo or vendor property we can point at, plus the map. Reported means the mid-2026 map — credible, single-source — is substantially all there is; those entries are attributed as such, carry verification flags in the copy, and should be read as leads, not reviews. The map’s author did the field a service; treating a census line as a product review would undo it.
The newcomer scorecard
The eight at a glance. Amber means the mid-2026 map is the main source — verify before you trust.
| Harness | Backer | The bet | Cost model | Evidence |
|---|---|---|---|---|
| Amp | Sourcegraph | Code-graph context, team scale | Paid, usage-based | Documented |
| Crush | Charm | Terminal craft as the product | Free + BYOK | Documented |
| OpenClaw | Independent | Iteration speed, broad tool surface | Unclear | Reported |
| Kilo CLI | Kilo (Cline→Roo lineage) | The fork that outlives vendors | Open source + BYOK | Documented lineage; CLI details reported |
| Hermes Agent | Unclear | Open-model-first agent | Unclear | Reported |
| DeepSeek-TUI | Unclear | Native at the V4 price floor | BYOK at floor prices | Reported |
| Ecodex | Independent | Calibration gates on actions | Unclear | Reported |
| Pi | Unclear | Minimalism | Unclear | Reported |
The eight, profiled
Amp (Sourcegraph)
The most prominent paid newcomer, per the mid-2026 map, and the least surprising bet on the list: Sourcegraph spent a decade indexing enterprise code, and Amp is that heritage pointed at the agent loop — context assembled from code intelligence rather than grep-and-hope. It ships terminal and editor surfaces, and its cultural stance is opinionated: reporting describes usage-based pricing from day one and a deliberate refusal to be free — the “someone pays for the tokens, on purpose” posture the die-off vindicated. Team features like shared threads and an opinionated model stance (Amp picks models rather than exposing a picker) recur in coverage.
The bet: the harness that already knows your codebase beats the harness that explores it. Best for: teams in large repos, especially existing Sourcegraph shops. Evidence: documented (vendor site + map).
Crush (Charm)
The best-documented entry here: Crush is Charm’s agent, from the company whose terminal-UI libraries half the modern CLI ecosystem is built on, carrying the codebase lineage of the OpenCode split forward. It is multi-model, MCP-capable, and LSP-aware, with session management and — being Charm — the best-looking agent session you can run; it also runs on Windows, which remains rarer than it should be in this category. Two cautions from the open-source harness roundup still apply: the agent loop is younger than the majors’, and check the license before assuming OSI-open — the lineage has not always been straightforwardly permissive.
The bet: the terminal is a designed surface, not a fallback. Best for: terminal-first BYOK developers; OpenCode alumni who want more polish. Evidence: documented (public repo).
OpenClaw
Per the mid-2026 map: a fast-moving independent entrant with unusual community momentum and a broad tool surface. That is genuinely about all that can be said with confidence, and the name adds a practical hazard — it has prior art in open source (a long-running game-engine reimplementation shares it), so confirm the repo you’re cloning is the harness you read about. Fast-moving independents are where this category’s surprises come from, in both directions: the die-off’s casualty list and its cult favorites were both born this way.
The bet: ship faster than the incumbents can copy. Best for: watchers, not daily-drivers, until documentation exists. Evidence: reported.
Kilo CLI
The survivor story in miniature. Cline forked into Roo Code; Roo Code shut down this year per mid-2026 reporting; Kilo forked from Roo, absorbed both lines’ features, and kept shipping — the fork-outlives-the-vendor pattern the die-off postmortem identified as the open lane’s structural advantage. The project’s editor-side lineage is well established; the map lists a terminal lane, extending it to CLI form. Open source, BYOK, provider-flexible.
The bet: continuity — your Cline-era muscle memory and configs, under governance that can’t be acquired out from under you. Best for: Cline and Roo refugees. Evidence: documented lineage; CLI specifics reported.
Hermes Agent
An entry on the map, a name that implies the open-model Hermes lineage, and very little else in the public record we can point at. Whether it is affiliated with the model family whose name it echoes is precisely the kind of thing to verify before granting anything a shell. File under watch.
The bet: unclear — plausibly an open-model-native agent. Best for: nobody yet. Evidence: reported.
DeepSeek-TUI
Per the map: a terminal entrant built around DeepSeek V4’s floor economics — and the economics are the substantiated part. V4 Flash lists at $0.14 per million input tokens and $0.28 output, the credible price floor for agentic coding, and a harness designed natively at that floor can afford habits frontier-priced tools ration: always-on background loops, generous re-reads, speculative parallel attempts. What’s unverified is nearly everything else, starting with whether it’s first-party or community-built — an enormous difference for trust, terms, and longevity. The data-residency and ToS questions that come with the Chinese CLI lane apply here unchanged.
The bet: design assumptions change when tokens cost almost nothing. Best for: cost-floor experimenters already comfortable in that lane. Evidence: reported.
Ecodex
The one carrying an actual new idea, and the reason this cohort is worth an article. Per the mid-2026 map, Ecodex builds “epistemic accountability” gates into the loop: the agent must state calibrated confidence in what it’s about to do, and actions it cannot justify above a threshold are blocked or escalated to the human. Every mature harness gates actions by tool class — write versus read, shell versus edit. Gating by the agent’s own epistemic state is a different axis, and a genuinely interesting one for anyone who has watched an agent proceed confidently off a wrong assumption.
The research behind the idea is real — language models can partially assess what they know, with measurable calibration (Language Models (Mostly) Know What They Know, arXiv) — and so is the catch: verbalized confidence degrades off-distribution and can be optimized into theater. A confidence gate is only as good as the calibration curve behind it, so the question for Ecodex is the reliability data, not the demo. Ask for it.
The bet: permissions keyed to justified confidence, not tool class. Best for: permission-model researchers and the safety-curious — watch and test, don’t daily-drive. Evidence: reported.
Pi
The thinnest entry on the map — a name, a minimalist reputation, and nothing else we can responsibly assert. The name is also maximally collision-prone: it is shared with a consumer AI assistant and at least one minimalist open-source agent project, so confirm you are installing the thing you researched. Minimalism is a legitimate bet in a category drowning in features; whether Pi actually makes it is unverifiable today.
The bet: less harness, on purpose. Best for: the curious with a sandbox. Evidence: reported.
Also on the radar, covered elsewhere: Kimi Code CLI rides the Chinese-plan economics story and gets its due in the Chinese CLI wave; Mistral’s Vibe CLI — the coding lane of the Le Chat–to–Vibe rebrand, per reporting — is a vendor tool, not a newcomer startup, and belongs in the vendor comparisons.
Genuinely new, or rebadged?
Strip the launch copy and ask what each newcomer re-engineered. Three ideas in this cohort pass the test.
Accountability gates (Ecodex). Permissioning by epistemic state is a new axis, not a new coat of paint — if the calibration holds up. It’s the only entry here that proposes changing what a harness is, which is why it gets the scrutiny above.
Terminal craft as the product (Crush). Dismissing UI polish as cosmetics misreads how operators actually work: session legibility, scannable diffs, and status you can parse at a glance are error-rate features when you supervise agents for hours. Charm is betting the surface is an organ, and Charm has the standing to make that argument.
Price-floor nativity (DeepSeek-TUI, if real as described). Harness design has always rationed tokens implicitly. A tool that assumes the $0.14 floor gets to spend context the way post-scarcity software spends RAM — a genuinely different design space, with the token-plan math to back why it matters.
The rest is execution rather than invention — which is not an insult. Kilo’s continuity bet and Amp’s code-graph context are serious engineering applied to known organs; the test, per the harness engineering framing and the awesome-harness-engineering literature it formalized, is which of the six subsystems a newcomer actually rebuilt. When the honest answer is “none, but the demo is pretty,” you’re looking at a rebadge, and the die-off already showed how those end.
Survivor criteria, applied
The die-off postmortem distilled four traits that separated 2026’s survivors from its casualties: a revenue model you can explain in one sentence, protocol-native plumbing, portable configs and data, and code that outlives the company. Score the newcomers against them and the table gets uncomfortable fast — which is the point.
| Newcomer | Explainable revenue | Protocol-native | Portable configs/data | Code outlives company |
|---|---|---|---|---|
| Amp | Yes — paid on purpose from day one | Reportedly yes | Test in trial — session export is the question | Weak — proprietary |
| Crush | Yes — BYOK, you pay inference | Yes — MCP, multi-provider | Yes — local sessions and configs | Mostly — source-available; read the license |
| OpenClaw | Unknown | Unknown | Unknown | Unknown |
| Kilo CLI | Yes — BYOK | Reportedly yes | Reportedly yes | Yes — open-source fork lineage |
| Hermes Agent | Unknown | Unknown | Unknown | Unknown |
| DeepSeek-TUI | BYOK at the floor — sustainable if community; strategic if first-party | Reportedly yes | Unknown | Unknown |
| Ecodex | Unknown | Unknown | Unknown | Unknown |
| Pi | Unknown | Unknown | Unknown | Unknown |
Read the unknowns correctly: they are not failing grades, they are unscored exams — and an unscored harness is one you sandbox, not one you daily-drive. On public evidence, Crush and Kilo CLI score strongest across the board; Amp scores where it counts most for longevity (someone pays, on purpose) and weakest on the trait that protects you if it dies anyway, which makes data portability the first thing to test in an Amp trial. The five reported-tier entries score “unknown” on nearly everything — the strongest possible argument for the protocol that closes this piece.
Who should try what
- Teams in monorepos, Sourcegraph shops, managers of shared agent practice: trial Amp first. Its bet aligns with your pain, and its pricing posture is a survival signal, not a bug.
- Terminal-first BYOK operators, OpenCode alumni, Windows developers underserved by the category: Crush, today. It’s the lowest-risk trial on the list — documented, inspectable, free to run on keys you already have.
- Cline and Roo refugees: Kilo CLI is your continuity path; verify the CLI lane’s maturity against the editor lineage before moving daily work.
- Cost-floor experimenters: DeepSeek-TUI, strictly sandboxed, strictly after confirming who builds it — with your existing V4 keys and the same trust posture you’d apply to any tool in that lane.
- Permission-model nerds and agent-safety people: watch Ecodex, and ask its builders for calibration data in public.
- Everyone else: OpenClaw, Hermes Agent, and Pi are watch-list entries. Bookmark, don’t install. Your daily-driver roster doesn’t have a gap these verifiably fill yet.
The safe-trial protocol
Every trait table above ends the same way: you learn the truth by trialing, and the die-off taught the cost of trialing carelessly. Here is the protocol — it assumes the worst and costs almost nothing when the tool turns out great.
Two weeks, four gates. The export path gets found on day zero, not on shutdown day.
1. Sandbox first. A new harness is unreviewed code you’re granting a shell — treat it that way. Container, VM, or a throwaway machine; a scoped API key with a hard spend cap; a clone of a real repo with no production credentials in reach. Reported-tier tools get the strictest version of this, because you can’t yet audit what you can’t yet find.
2. Find the export path before day one. Where do sessions live on disk, in what format? Do configs follow AGENTS.md-style conventions you can carry out? Can you get transcripts and instruction files back without the vendor’s help? This is the die-off’s exit checklist run in reverse — as an entrance exam. A harness that fails it can still be worth trialing; it can’t be worth trusting.
3. Run the two-week bake-off. Week one, shadow mode: replay tasks you already completed with your daily driver, where you know what good looks like, and compare. Week two, live mode: real tasks, with the incumbent on standby and a log of every intervention. Score four things — completion rate, interventions per task, tokens burned, and wall-clock to merged. Gut feel lies about new tools in both directions; the log doesn’t.
4. Decide at the gate you set in advance. Write the keep/drop criteria before the trial starts, decide on day fourteen, and export your sessions either way — the trial’s transcripts are evidence you paid for, whatever the verdict.
Product note: Trialing newcomers multiplies exactly the sprawl Automater Lite exists for. Any CLI that writes transcripts joins its local archive — the roster covers 10+ providers — so a two-week bake-off adds searchable history instead of orphaning it, and local per-provider token metering tells you what the trial actually burned before you commit. Free, on automater.ai.
Run it that way and the fleet you already manage absorbs a newcomer without drama: worst case, you spent two sandboxed weeks and kept the transcripts; best case, one of the eight earns a permanent slot. Some of this cohort will headline next year’s casualty list — that’s not cynicism, it’s the base rate — and the operators who trial with export paths and evidence tiers will be the ones for whom it never matters.
FAQ: new AI coding tools in 2026
What are the best new AI coding tools in 2026?
Among post-die-off newcomers, Amp (Sourcegraph’s paid, code-graph-driven harness) and Crush (Charm’s open, terminal-craft agent) are the most substantiated as of August 2026. OpenClaw, Kilo CLI, Hermes Agent, DeepSeek-TUI, Ecodex, and Pi appear in credible mid-2026 reporting but need verification before daily use.
What is Amp by Sourcegraph?
Amp is Sourcegraph’s agentic coding tool, per mid-2026 reporting the most prominent of the new paid CLIs. Its bet is context: assembling what the agent sees from Sourcegraph’s code-intelligence heritage rather than ad-hoc search, with team features and usage-based pricing — someone pays for the tokens on purpose.
What is Crush CLI?
Crush is Charm’s open, multi-model coding agent for the terminal — MCP-capable, LSP-aware, session-based, and built with the terminal-UI craft Charm’s libraries are known for, Windows included. It descends from the OpenCode split; check its source-available license before assuming OSI-open, and bring your own API keys.
What are calibration gates in AI coding agents?
Calibration gates — the “epistemic accountability” idea attributed to Ecodex in mid-2026 reporting — make an agent state calibrated confidence before acting, blocking or escalating actions it can’t justify above a threshold. Research shows models partially know what they know; whether that calibration survives real coding work is the open question.
How do I safely try a new AI coding CLI?
Sandbox it like unreviewed code: container or VM, scoped API key with a spend cap, no production credentials. Find the session export path before day one. Run a two-week bake-off — one shadow week on solved tasks, one live week with your incumbent on standby — and decide against criteria you wrote in advance.
Sources
- Coding CLIs in mid-2026: the engineer’s map (dev.to)
- Gemini CLI shutdown takes effect as CI/CD pipelines break (TechTimes)
- Crush — Charmbracelet (GitHub)
- Amp — Sourcegraph (ampcode.com)
- Kilo Code (kilocode.ai)
- Best open-source coding models 2026 (morphllm.com)
- Language Models (Mostly) Know What They Know (arXiv)
- awesome-harness-engineering (GitHub)
