Write a Finish Line the Transcript Can Prove
An agent goal completion condition is judged from the transcript alone. Write an end state, a stated check, constraints and a bound; back it with a Stop hook.
Go deeper. Build your own.
The judge that decides whether your overnight goal is done never runs a command. Claude Code’s /goal hands the verdict to a small fast model after every turn, and the goal docs are plain about its reach: “It doesn’t run commands or read files independently, so write the condition as something Claude’s own output can demonstrate.” Write “the migration is complete” and you have asked that judge to take the worker’s word for it.
You can write an agent goal completion condition that holds up. By Tuesday, every goal you hand a lane has four parts: an observable end state, a stated check whose output lands in the transcript, constraints on what must not change, and a turn or time bound. Where the transcript can be gamed, a Stop-hook script runs the check itself. And a run that ends on a narrated plan or a list of next steps for you is logged as not done, whatever the judge said.
Chatbots suggest; agents act, and an agent in goal mode keeps acting until something tells it to stop. The finish line is that something.
April to September: three vendors ship goal loops, and one documents its judge
Claude Code added /goal in 2.1.139 on May 11, 2026, per the changelog. The mechanism is one sentence in the docs: “After each turn, a small fast model checks whether the condition holds.” It returns Not yet met, Met or Impossible, each with a short reason; the transcript shows every verdict, and Ctrl+O shows the reason.
The evaluator “defaults to Haiku on the Claude API”; on a third-party provider, the docs send you to your provider’s page, and ANTHROPIC_DEFAULT_HAIKU_MODEL moves it along with every other small-fast-model job. Completion “is decided by a fresh model rather than the one doing the work.”
Screenshot: Claude Code Docs, “Keep Claude working toward a goal” (undated page), captured Sep 28, 2026.
The same page shows the right instinct. A condition that all tests in test/auth pass “works because Claude runs the tests and the result appears in the transcript for the evaluator to read.” A practitioner explainer from May put the shape well: two models in a loop, one working and one reading.
Codex took a different road. Goal mode shipped as an experiment in Codex CLI 0.128.0 on Apr 30, 2026, and the Codex changelog entry dated May 21 (app 26.519) took it out of experimental: “Goal mode is no longer an experimental feature and is available in the Codex app, IDE extension, and CLI.”
The docs describe no separate judge. The long-running work guide says “Write a goal that lets ChatGPT verify its own progress”, and asks for an outcome, constraints and verification. On Sep 17, CLI 0.155.0 added a stall guard through PR #44320: “Stop this loop by marking the goal as blocked after three consecutive empty turns with no other activity.”
Kimi Code’s 0.43.0 changelog (Sep 14) loosened its time budgets: “Goal time budgets no longer count time spent with the session closed, and the 24-hour limit is removed.” Its goal mode guide has the best one-line definition of the feature: “a normal prompt says what to do next, a goal says what must become true.”
And on Sep 22, Anthropic’s claude.dev guide to Opus 5.5 compressed the whole practice: “Give the whole task in one message. Name the finish line, like “the tests pass” or “every endpoint is migrated.” Then let it cook.”
Why an agent goal completion condition has to be provable from the transcript
A judge that reads only the conversation grades what was said. The goal docs don’t say the evaluator can tell real command output from text the worker merely typed; that the transcript can be gamed is our inference from “It doesn’t run commands or read files independently”, not documented behavior. Nobody needs a malicious model for it to bite. A worker that writes “all tests pass” after running a subset, or summarizes a check it meant to run, gives the judge the same words a finished run would.
So the condition has to ask for evidence whose absence is visible: a named command, its output and its exit code, in this session. The rest of this runbook is how to write that, and what to do where words aren’t enough.
Step 1: Give every agent goal completion condition four parts
The Claude Code goal docs name the parts; Codex’s guide asks for the same under other names (an outcome, constraints, and verification: “Add tests, measurements, or review criteria that prove the work is complete”). The bound is the part people forget.
| Part | What it pins down | Good | Bad, and why the judge can’t use it |
|---|---|---|---|
| End state | What is true when the work is done, observable from outside | “Every handler in src/payments/ calls the new client, and the old client file is deleted” |
“The payment migration is finished”: a claim, not a state |
| Stated check | The command that proves it, run in this session, output in the transcript | “npm test exits 0” or “git status is clean” (the docs’ examples) |
“Tests should pass”: asks for no run, so a sentence satisfies it |
| Constraints | What must not change on the way there | “No other test file is modified” (the docs’ example) | “Be careful with the tests”: nothing to check |
| Bound | When to stop trying | “or stop after 20 turns” (the docs’ example) | None: the goal docs set no default turn cap |
End state. Make it a state, not an activity. “Every handler calls the new client” can be checked; “migrate the handlers” can only be attempted. Make it countable where you can: files that must exist or be gone, a count that must reach zero, a queue that must be empty.
Stated check. This part carries the weight. Name the exact command and require its exit code in the transcript. Our practice is to ask for the summary line and the exit code printed explicitly, because a wall of log output buries the one line the judge needs.
Constraints. This is where goal runs cheat without meaning to. The fastest way to make a failing suite pass is to change the suite, so every goal that names a test check also names the tests that must not change, and any config that must stay put.
Bound. From the docs: “To bound how long a goal runs, include a turn or time clause in the condition, such as or stop after 20 turns. Claude reports progress against that clause each turn and the evaluator judges it from the conversation.”
Step 2: Turn a good prompt into a condition the judge can grade
The Opus 5.5 guide’s own finish-line example is a prompt, not a goal: “Migrate the payment endpoints from the old client to the new one. Done means: every endpoint uses the new client, the old client is deleted, and the test suite passes. Stop and ask me only if a test fails for a reason you can’t explain.” It names the end state well. As a /goal condition it still needs a stated check, a constraint and a bound.
Illustrative conditions, each well under the 4,000-character limit, with made-up paths:
/goal Every handler under src/payments/ imports PayClientV2, and
src/legacy/pay-client.ts is deleted. Prove it in this session: run
`grep -rn "legacy/pay-client" src/` and show that it prints nothing, then
run `npm test` and show the summary line and exit code 0. Constraints: no
file under test/ outside test/payments/ is modified; package.json is
unchanged. Or stop after 25 turns and report what remains.
/goal The lint job is green: `npm run lint` exits 0, with its last lines
shown. Only files that lint reported are edited, and no lint rule is
disabled or ignored. Or stop after 15 turns.
/goal test/queue/retry.test.ts is no longer flaky: it passes 10 runs in a
row, executed as one loop in this session with each exit code printed. The
test's assertions are unchanged. Or stop after 20 turns.
Each line of those conditions answers a question the judge would otherwise have to guess. The grep proves absence, which a summary can’t. The constraint on test/ closes the cheapest shortcut, and the package.json constraint stops a dependency swap from “fixing” the build. The loop count turns “no longer flaky” from an opinion into a number.
Two conditions to refuse on sight: anything whose end state is a feeling (“the code is clean”, “refactor until it’s good”) and anything whose check is the worker’s own report (“confirm that everything works”). Both pass the moment the worker says they do.
The four parts port. Codex’s guide reads the goal as the task’s own test (“The goal text becomes both the first prompt and the completion criteria for the task”), and Kimi’s wants the objective “naming the finish line and the stop condition”. Write the condition once in your task template and it works in all three.
Step 3: Bound the run yourself, because the defaults mostly bound waiting
The clause in your condition is the only turn cap /goal has; the docs set none by default. The no-progress guard, which stops the loop when Claude keeps answering the evaluator without using tools “for several turns in a row”, publishes no number. The other limits govern waiting: while background subagents or shells run, evaluation defers and the goal checks in at 30 minutes, then 1 hour, then every 2 hours (CLAUDE_CODE_GOAL_CHECKIN_MINUTES, with 0 turning check-ins off), at most three idle check-ins per goal between your prompts (v2.1.246+).
Our rule of thumb: set the turn clause at about twice what a clean run of a similar task took, and add a time clause for anything that waits on CI. And remember that “A goal doesn’t change your permission mode.” In manual mode an overnight goal simply waits at the first tool call your settings don’t allow; spotting a lane that is quietly waiting is the job of stall flags.
Each vendor guards the loop differently:
Claude Code /goal |
Codex Goal mode | Kimi Code goal mode | |
|---|---|---|---|
| Status | Since 2.1.139 (May 11, 2026) | Experimental in CLI 0.128.0 (Apr 30); out of experimental May 21 (app 26.519) | Time-budget change in 0.43.0 (Sep 14) |
| Who judges done | A separate small fast model after each turn; Haiku on the Claude API | No separate judge documented; the goal “lets ChatGPT verify its own progress” | The guide doesn’t say |
| Condition length | Up to 4,000 characters | Not in the docs we read | Up to 4,000 characters |
| Stall guard | No-progress guard (count unpublished); three idle check-ins | Blocked after three empty continuation turns (0.155.0, Sep 17) | Ends complete, paused or blocked |
| Bound | Your turn or time clause | Pauses when it needs a decision | Time budgets; 24-hour limit removed in 0.43.0 |
| Headless | Works in -p |
In the CLI; headless use not in the docs we read | kimi -p "/goal ..." exits 0 complete, 3 blocked, 6 paused |
| Permissions | Unchanged by the goal | “keeps the same sandbox and approval policy” | Not in the docs we read |
Every number is from vendor docs, read Sep 28, 2026. The last tile is the one to act on.
Kimi’s headless exit codes are worth copying as a convention even where a CLI doesn’t give you them: complete, blocked and paused are three different mornings, and a scheduler that sees only “exited” treats them the same.
Step 4: Back the transcript with a command Stop hook that runs the check
Where a wrong Met is expensive, don’t make the judge rely on the worker’s typing. Run the check yourself on every stop.
Start with what /goal already is: “/goal is a wrapper around a session-scoped prompt-based Stop hook.” A prompt hook is one model call that sees only the hook input, so “back it with a Stop hook” has to mean a different kind. The hooks guide draws the line: “Use prompt hooks when the hook input data alone is enough to make a decision. Use agent hooks when you need to verify something against the actual state of the codebase.” Then it warns: “Agent hooks are experimental. Behavior and configuration may change in future releases. For production workflows, prefer command hooks.”
Screenshot: Claude Code Docs, “Automate actions with hooks” (undated page), captured Sep 28, 2026.
The evaluator grades the conversation. The script grades the repo.
A command hook is a script. Per the hooks reference, a Stop hook receives stop_hook_active, last_assistant_message, background_tasks and session_crons, and exiting with code 2 and a message on stderr (or returning decision: "block" with a reason) keeps Claude working. The settings below mirror the docs’ Stop-hook examples with the type switched to command; the path and script are illustrative:
{
"hooks": {
"Stop": [
{
"hooks": [
{ "type": "command", "command": ".claude/hooks/finish-line.sh" }
]
}
]
}
}
#!/usr/bin/env bash
# .claude/hooks/finish-line.sh (illustrative): check the repo on every stop
# instead of trusting what the transcript says.
set -uo pipefail
cat > /dev/null # the hook input arrives on stdin; this check ignores it
if grep -rqn "legacy/pay-client" src/; then
echo "Not done: src/ still references legacy/pay-client." >&2
exit 2
fi
if ! npm test --silent > /tmp/finish-line.log 2>&1; then
echo "Not done: npm test failed. Last lines:" >&2
tail -n 15 /tmp/finish-line.log >&2
exit 2
fi
outside=$(git diff --name-only main -- test/ | grep -v '^test/payments/' || true)
if [ -n "$outside" ]; then
echo "Constraint broken, test files changed: $outside" >&2
exit 2
fi
exit 0
Know its limits before you trust it. “Claude Code applies an 8-consecutive-continuation cap: after stop hooks have continued the turn eight times in a row, Claude Code overrides the next block and ends the turn.” At the default, a run that fails the check a ninth time ends its turn anyway (CLAUDE_CODE_STOP_HOOK_BLOCK_CAP raises the cap), so keep the hook’s own log as the record.
The docs don’t say what happens when /goal judges Met while your command hook blocks in the same turn: test that on a toy repo before a real run depends on it. And /goal is unavailable when disableAllHooks is true or allowManagedHooksOnly is set, so check both on every machine that runs goals.
A hook is code that runs on every stop, so review and register it like any other; hooks that arrive unauthored are the other side of this pattern. Behind the hook sits the gate it can’t replace, CI and review at merge, covered in overnight merge gates.
Step 5: Log narrated plans and next-steps endings as not done
The Opus 5.5 guide names the endings to reject: “a summary that names the next step without taking it, an offer to continue, or a list of choices that don’t block the work.” Under a goal, those endings are how a bounded run spends its last turns sounding finished.
Grade every goal run in the morning with one rule: done means the stated check’s output is in the final turns and the hook’s log agrees. Everything else is not done, whatever the verdict said.
- A plan or summary without the check output: not done; re-queue with the same condition.
- “Next steps for you” or a handoff list: not done; the items become the next condition or a question for the owner.
- Met, but the hook’s log shows a failed check: not done, and a bug against the condition, because transcript and repo disagree.
- Impossible, with a reason: a legitimate stop that finished nothing; read the reason before re-queuing.
Count not-done runs per lane per week. A lane whose Met verdicts keep failing the hook has a condition-writing problem, not a model problem. Put the grading rule where every CLI reads it: the AGENTS.md run contract carries the clause, and TASKS.md receipts carries the evidence each tick needs.
How a goal reads Met while the work isn’t done
Typed, not run. The worker writes the result of a check it never ran. Signal: check output with no tool call before it, or a hook that disagrees. Fix: require the output, keep the hook.
The narrowed check. One file’s tests run; the suite gets reported. Signal: the test count in the output drops against the last full run. Fix: name the exact command, with no arguments left to the worker.
The edited test. Signal: diffs under test/ outside the allowed path. Fix: the constraint clause, plus the hook’s diff check.
The thrash. Each turn fails the same check a new way. Signal: repeated blocks with the same reason, ending at the cap of eight. Fix: a lower turn bound and the loop limits in CI fix-loop guards.
The moved judge. A provider switch changes the evaluator, and ANTHROPIC_DEFAULT_HAIKU_MODEL moves it along with the haiku alias and background summarization. Signal: verdict reasons change in strictness after a config change. Fix: re-run one known-good and one known-bad goal after any change.
No hooks, no goal. Signal: /goal is unavailable on a machine with disableAllHooks or allowManagedHooksOnly set. Fix: know which machines those are before you schedule goals on them.
Keep the judge outside the worker
A goal loop is a worker and a judge. Claude Code ships both halves, Codex asks the worker to verify its own progress, and either way the judge is the half to own. The evaluator grades the transcript, your hook grades the repo, and your morning review grades the hook’s log.
Keep those records with the rest of the fleet’s evidence, so a disputed Met can be replayed turn by turn, which is the job fleet replay exists for. No vendor’s judge sees the other vendors’ lanes, and a condition written once should be graded the same way in all of them.
FAQ
What does Claude Code’s /goal evaluator check?
After each turn, a small fast model reads the conversation and returns Not yet met, Met or Impossible, each with a short reason. It defaults to Haiku on the Claude API, calls no tools and reads no files, so it can only judge what Claude’s own output has already shown in the transcript.
How do I stop an agent goal from running forever?
Write the bound into the condition, such as “or stop after 20 turns”, because Claude Code’s /goal sets no default turn cap. Claude reports progress against that clause each turn. Codex blocks a goal after three empty continuation turns, and Kimi Code uses time budgets, so check each vendor’s guard.
Sources
- Claude Code: Keep Claude working toward a goal — the evaluator, conditions, check-ins
- Claude Code hooks guide — prompt, agent and command Stop hooks
- Claude Code hooks reference — Stop input, exit code 2, the 8-block cap
- Claude Code changelog —
/goalin 2.1.139, May 11, 2026 - Codex changelog — Goal mode out of experimental, May 21, 2026; CLI 0.155.0
- Codex: long-running work — goals as completion criteria
- openai/codex PR #44320 — block goals after three empty turns
- Kimi Code CLI changelog — 0.43.0, Sep 14, 2026
- Kimi Code goal mode guide — exit codes 0, 3, 6
- Anthropic claude.dev Blog: “Getting the most out of Opus 5.5 in Claude and Claude Code” — Addy Osmani, Sep 22, 2026
