Write a Finish Line the Transcript Can Prove

An agent goal completion condition is judged from the transcript alone. Write an end state, a stated check, constraints and a bound; back it with a Stop hook.

Agent goal completion condition proven in the transcript: a terminal card showing npm test with exit 0 and a clean git status above a checkered finish lineAgent goal completion condition proven in the transcript: a terminal card showing npm test with exit 0 and a clean git status above a checkered finish line
The evaluator can only grade what the transcript shows. Put the proof there.

The judge that decides whether your overnight goal is done never runs a command. Claude Code’s /goal hands the verdict to a small fast model after every turn, and the goal docs are plain about its reach: “It doesn’t run commands or read files independently, so write the condition as something Claude’s own output can demonstrate.” Write “the migration is complete” and you have asked that judge to take the worker’s word for it.

You can write an agent goal completion condition that holds up. By Tuesday, every goal you hand a lane has four parts: an observable end state, a stated check whose output lands in the transcript, constraints on what must not change, and a turn or time bound. Where the transcript can be gamed, a Stop-hook script runs the check itself. And a run that ends on a narrated plan or a list of next steps for you is logged as not done, whatever the judge said.

Chatbots suggest; agents act, and an agent in goal mode keeps acting until something tells it to stop. The finish line is that something.

April to September: three vendors ship goal loops, and one documents its judge

Claude Code added /goal in 2.1.139 on May 11, 2026, per the changelog. The mechanism is one sentence in the docs: “After each turn, a small fast model checks whether the condition holds.” It returns Not yet met, Met or Impossible, each with a short reason; the transcript shows every verdict, and Ctrl+O shows the reason.

The evaluator “defaults to Haiku on the Claude API”; on a third-party provider, the docs send you to your provider’s page, and ANTHROPIC_DEFAULT_HAIKU_MODEL moves it along with every other small-fast-model job. Completion “is decided by a fresh model rather than the one doing the work.”

Claude Code Docs page on goals, section Write an effective condition, explaining that the evaluator judges the condition against what Claude has surfaced in the conversation and does not run commands or read files, followed by a list that begins with one measurable end state and a stated check Screenshot: Claude Code Docs, “Keep Claude working toward a goal” (undated page), captured Sep 28, 2026.

The same page shows the right instinct. A condition that all tests in test/auth pass “works because Claude runs the tests and the result appears in the transcript for the evaluator to read.” A practitioner explainer from May put the shape well: two models in a loop, one working and one reading.

Codex took a different road. Goal mode shipped as an experiment in Codex CLI 0.128.0 on Apr 30, 2026, and the Codex changelog entry dated May 21 (app 26.519) took it out of experimental: “Goal mode is no longer an experimental feature and is available in the Codex app, IDE extension, and CLI.”

The docs describe no separate judge. The long-running work guide says “Write a goal that lets ChatGPT verify its own progress”, and asks for an outcome, constraints and verification. On Sep 17, CLI 0.155.0 added a stall guard through PR #44320: “Stop this loop by marking the goal as blocked after three consecutive empty turns with no other activity.”

Kimi Code’s 0.43.0 changelog (Sep 14) loosened its time budgets: “Goal time budgets no longer count time spent with the session closed, and the 24-hour limit is removed.” Its goal mode guide has the best one-line definition of the feature: “a normal prompt says what to do next, a goal says what must become true.”

And on Sep 22, Anthropic’s claude.dev guide to Opus 5.5 compressed the whole practice: “Give the whole task in one message. Name the finish line, like “the tests pass” or “every endpoint is migrated.” Then let it cook.”

Why an agent goal completion condition has to be provable from the transcript

A judge that reads only the conversation grades what was said. The goal docs don’t say the evaluator can tell real command output from text the worker merely typed; that the transcript can be gamed is our inference from “It doesn’t run commands or read files independently”, not documented behavior. Nobody needs a malicious model for it to bite. A worker that writes “all tests pass” after running a subset, or summarizes a check it meant to run, gives the judge the same words a finished run would.

So the condition has to ask for evidence whose absence is visible: a named command, its output and its exit code, in this session. The rest of this runbook is how to write that, and what to do where words aren’t enough.

Step 1: Give every agent goal completion condition four parts

The Claude Code goal docs name the parts; Codex’s guide asks for the same under other names (an outcome, constraints, and verification: “Add tests, measurements, or review criteria that prove the work is complete”). The bound is the part people forget.

Part What it pins down Good Bad, and why the judge can’t use it
End state What is true when the work is done, observable from outside “Every handler in src/payments/ calls the new client, and the old client file is deleted” “The payment migration is finished”: a claim, not a state
Stated check The command that proves it, run in this session, output in the transcript “npm test exits 0” or “git status is clean” (the docs’ examples) “Tests should pass”: asks for no run, so a sentence satisfies it
Constraints What must not change on the way there “No other test file is modified” (the docs’ example) “Be careful with the tests”: nothing to check
Bound When to stop trying “or stop after 20 turns” (the docs’ example) None: the goal docs set no default turn cap

End state. Make it a state, not an activity. “Every handler calls the new client” can be checked; “migrate the handlers” can only be attempted. Make it countable where you can: files that must exist or be gone, a count that must reach zero, a queue that must be empty.

Stated check. This part carries the weight. Name the exact command and require its exit code in the transcript. Our practice is to ask for the summary line and the exit code printed explicitly, because a wall of log output buries the one line the judge needs.

Constraints. This is where goal runs cheat without meaning to. The fastest way to make a failing suite pass is to change the suite, so every goal that names a test check also names the tests that must not change, and any config that must stay put.

Bound. From the docs: “To bound how long a goal runs, include a turn or time clause in the condition, such as or stop after 20 turns. Claude reports progress against that clause each turn and the evaluator judges it from the conversation.”

Step 2: Turn a good prompt into a condition the judge can grade

The Opus 5.5 guide’s own finish-line example is a prompt, not a goal: “Migrate the payment endpoints from the old client to the new one. Done means: every endpoint uses the new client, the old client is deleted, and the test suite passes. Stop and ask me only if a test fails for a reason you can’t explain.” It names the end state well. As a /goal condition it still needs a stated check, a constraint and a bound.

Illustrative conditions, each well under the 4,000-character limit, with made-up paths:

/goal Every handler under src/payments/ imports PayClientV2, and
src/legacy/pay-client.ts is deleted. Prove it in this session: run
`grep -rn "legacy/pay-client" src/` and show that it prints nothing, then
run `npm test` and show the summary line and exit code 0. Constraints: no
file under test/ outside test/payments/ is modified; package.json is
unchanged. Or stop after 25 turns and report what remains.

/goal The lint job is green: `npm run lint` exits 0, with its last lines
shown. Only files that lint reported are edited, and no lint rule is
disabled or ignored. Or stop after 15 turns.

/goal test/queue/retry.test.ts is no longer flaky: it passes 10 runs in a
row, executed as one loop in this session with each exit code printed. The
test's assertions are unchanged. Or stop after 20 turns.

Each line of those conditions answers a question the judge would otherwise have to guess. The grep proves absence, which a summary can’t. The constraint on test/ closes the cheapest shortcut, and the package.json constraint stops a dependency swap from “fixing” the build. The loop count turns “no longer flaky” from an opinion into a number.

Two conditions to refuse on sight: anything whose end state is a feeling (“the code is clean”, “refactor until it’s good”) and anything whose check is the worker’s own report (“confirm that everything works”). Both pass the moment the worker says they do.

The four parts port. Codex’s guide reads the goal as the task’s own test (“The goal text becomes both the first prompt and the completion criteria for the task”), and Kimi’s wants the objective “naming the finish line and the stop condition”. Write the condition once in your task template and it works in all three.

Step 3: Bound the run yourself, because the defaults mostly bound waiting

The clause in your condition is the only turn cap /goal has; the docs set none by default. The no-progress guard, which stops the loop when Claude keeps answering the evaluator without using tools “for several turns in a row”, publishes no number. The other limits govern waiting: while background subagents or shells run, evaluation defers and the goal checks in at 30 minutes, then 1 hour, then every 2 hours (CLAUDE_CODE_GOAL_CHECKIN_MINUTES, with 0 turning check-ins off), at most three idle check-ins per goal between your prompts (v2.1.246+).

Our rule of thumb: set the turn clause at about twice what a clean run of a similar task took, and add a time clause for anything that waits on CI. And remember that “A goal doesn’t change your permission mode.” In manual mode an overnight goal simply waits at the first tool call your settings don’t allow; spotting a lane that is quietly waiting is the job of stall flags.

Each vendor guards the loop differently:

Claude Code /goal Codex Goal mode Kimi Code goal mode
Status Since 2.1.139 (May 11, 2026) Experimental in CLI 0.128.0 (Apr 30); out of experimental May 21 (app 26.519) Time-budget change in 0.43.0 (Sep 14)
Who judges done A separate small fast model after each turn; Haiku on the Claude API No separate judge documented; the goal “lets ChatGPT verify its own progress” The guide doesn’t say
Condition length Up to 4,000 characters Not in the docs we read Up to 4,000 characters
Stall guard No-progress guard (count unpublished); three idle check-ins Blocked after three empty continuation turns (0.155.0, Sep 17) Ends complete, paused or blocked
Bound Your turn or time clause Pauses when it needs a decision Time budgets; 24-hour limit removed in 0.43.0
Headless Works in -p In the CLI; headless use not in the docs we read kimi -p "/goal ..." exits 0 complete, 3 blocked, 6 paused
Permissions Unchanged by the goal “keeps the same sandbox and approval policy” Not in the docs we read

Stat tiles for the numbers that bound an agent goal loop: a 4,000-character condition limit, a 2-hour steady check-in interval after 30 minutes and 1 hour, 3 idle check-ins per goal, 8 Stop-hook blocks in a row before an override, 3 empty turns before Codex blocks a goal, and no default turn cap in /goalStat tiles for the numbers that bound an agent goal loop: a 4,000-character condition limit, a 2-hour steady check-in interval after 30 minutes and 1 hour, 3 idle check-ins per goal, 8 Stop-hook blocks in a row before an override, 3 empty turns before Codex blocks a goal, and no default turn cap in /goal Every number is from vendor docs, read Sep 28, 2026. The last tile is the one to act on.

Kimi’s headless exit codes are worth copying as a convention even where a CLI doesn’t give you them: complete, blocked and paused are three different mornings, and a scheduler that sees only “exited” treats them the same.

Step 4: Back the transcript with a command Stop hook that runs the check

Where a wrong Met is expensive, don’t make the judge rely on the worker’s typing. Run the check yourself on every stop.

Start with what /goal already is: “/goal is a wrapper around a session-scoped prompt-based Stop hook.” A prompt hook is one model call that sees only the hook input, so “back it with a Stop hook” has to mean a different kind. The hooks guide draws the line: “Use prompt hooks when the hook input data alone is enough to make a decision. Use agent hooks when you need to verify something against the actual state of the codebase.” Then it warns: “Agent hooks are experimental. Behavior and configuration may change in future releases. For production workflows, prefer command hooks.”

Claude Code Docs hooks guide showing an agent-type Stop hook configuration whose prompt asks it to verify that all unit tests pass, followed by the guidance to use prompt hooks when hook input alone is enough and agent hooks to verify against the codebase Screenshot: Claude Code Docs, “Automate actions with hooks” (undated page), captured Sep 28, 2026.

Diagram of the goal loop: a turn feeds the transcript, the evaluator reads only the transcript and returns met, impossible or not yet met, while a command Stop hook runs the check against the repo and blocks the stop with exit code 2Diagram of the goal loop: a turn feeds the transcript, the evaluator reads only the transcript and returns met, impossible or not yet met, while a command Stop hook runs the check against the repo and blocks the stop with exit code 2 The evaluator grades the conversation. The script grades the repo.

A command hook is a script. Per the hooks reference, a Stop hook receives stop_hook_active, last_assistant_message, background_tasks and session_crons, and exiting with code 2 and a message on stderr (or returning decision: "block" with a reason) keeps Claude working. The settings below mirror the docs’ Stop-hook examples with the type switched to command; the path and script are illustrative:

{
  "hooks": {
    "Stop": [
      {
        "hooks": [
          { "type": "command", "command": ".claude/hooks/finish-line.sh" }
        ]
      }
    ]
  }
}
#!/usr/bin/env bash
# .claude/hooks/finish-line.sh (illustrative): check the repo on every stop
# instead of trusting what the transcript says.
set -uo pipefail
cat > /dev/null   # the hook input arrives on stdin; this check ignores it

if grep -rqn "legacy/pay-client" src/; then
  echo "Not done: src/ still references legacy/pay-client." >&2
  exit 2
fi
if ! npm test --silent > /tmp/finish-line.log 2>&1; then
  echo "Not done: npm test failed. Last lines:" >&2
  tail -n 15 /tmp/finish-line.log >&2
  exit 2
fi
outside=$(git diff --name-only main -- test/ | grep -v '^test/payments/' || true)
if [ -n "$outside" ]; then
  echo "Constraint broken, test files changed: $outside" >&2
  exit 2
fi
exit 0

Know its limits before you trust it. “Claude Code applies an 8-consecutive-continuation cap: after stop hooks have continued the turn eight times in a row, Claude Code overrides the next block and ends the turn.” At the default, a run that fails the check a ninth time ends its turn anyway (CLAUDE_CODE_STOP_HOOK_BLOCK_CAP raises the cap), so keep the hook’s own log as the record.

The docs don’t say what happens when /goal judges Met while your command hook blocks in the same turn: test that on a toy repo before a real run depends on it. And /goal is unavailable when disableAllHooks is true or allowManagedHooksOnly is set, so check both on every machine that runs goals.

A hook is code that runs on every stop, so review and register it like any other; hooks that arrive unauthored are the other side of this pattern. Behind the hook sits the gate it can’t replace, CI and review at merge, covered in overnight merge gates.

Step 5: Log narrated plans and next-steps endings as not done

The Opus 5.5 guide names the endings to reject: “a summary that names the next step without taking it, an offer to continue, or a list of choices that don’t block the work.” Under a goal, those endings are how a bounded run spends its last turns sounding finished.

Grade every goal run in the morning with one rule: done means the stated check’s output is in the final turns and the hook’s log agrees. Everything else is not done, whatever the verdict said.

  • A plan or summary without the check output: not done; re-queue with the same condition.
  • “Next steps for you” or a handoff list: not done; the items become the next condition or a question for the owner.
  • Met, but the hook’s log shows a failed check: not done, and a bug against the condition, because transcript and repo disagree.
  • Impossible, with a reason: a legitimate stop that finished nothing; read the reason before re-queuing.

Count not-done runs per lane per week. A lane whose Met verdicts keep failing the hook has a condition-writing problem, not a model problem. Put the grading rule where every CLI reads it: the AGENTS.md run contract carries the clause, and TASKS.md receipts carries the evidence each tick needs.

How a goal reads Met while the work isn’t done

Typed, not run. The worker writes the result of a check it never ran. Signal: check output with no tool call before it, or a hook that disagrees. Fix: require the output, keep the hook.

The narrowed check. One file’s tests run; the suite gets reported. Signal: the test count in the output drops against the last full run. Fix: name the exact command, with no arguments left to the worker.

The edited test. Signal: diffs under test/ outside the allowed path. Fix: the constraint clause, plus the hook’s diff check.

The thrash. Each turn fails the same check a new way. Signal: repeated blocks with the same reason, ending at the cap of eight. Fix: a lower turn bound and the loop limits in CI fix-loop guards.

The moved judge. A provider switch changes the evaluator, and ANTHROPIC_DEFAULT_HAIKU_MODEL moves it along with the haiku alias and background summarization. Signal: verdict reasons change in strictness after a config change. Fix: re-run one known-good and one known-bad goal after any change.

No hooks, no goal. Signal: /goal is unavailable on a machine with disableAllHooks or allowManagedHooksOnly set. Fix: know which machines those are before you schedule goals on them.

Keep the judge outside the worker

A goal loop is a worker and a judge. Claude Code ships both halves, Codex asks the worker to verify its own progress, and either way the judge is the half to own. The evaluator grades the transcript, your hook grades the repo, and your morning review grades the hook’s log.

Keep those records with the rest of the fleet’s evidence, so a disputed Met can be replayed turn by turn, which is the job fleet replay exists for. No vendor’s judge sees the other vendors’ lanes, and a condition written once should be graded the same way in all of them.

FAQ

What does Claude Code’s /goal evaluator check?

After each turn, a small fast model reads the conversation and returns Not yet met, Met or Impossible, each with a short reason. It defaults to Haiku on the Claude API, calls no tools and reads no files, so it can only judge what Claude’s own output has already shown in the transcript.

How do I stop an agent goal from running forever?

Write the bound into the condition, such as “or stop after 20 turns”, because Claude Code’s /goal sets no default turn cap. Claude reports progress against that clause each turn. Codex blocks a goal after three empty continuation turns, and Kimi Code uses time budgets, so check each vendor’s guard.

Sources

YOU'RE THROUGH THIS ONE.

Keep connecting the dots.

Back to the library