TASKS.md With Receipts: Every Tick Needs Evidence
A TASKS.md checklist survives compaction, but a tick is only a claim. Give every tick a receipt a verifier can re-run, and revert any tick that has none.
Go deeper. Build your own.
A tick costs one keystroke, and - [x] services/billing: fixed reads the same whether the fix ran, half ran or was only described. That is the gap in the advice Anthropic gave Opus 5.5 users on Sep 22.
For long runs it suggests a one-line prompt: “Keep a checklist in TASKS.md. Tick each item when it’s done, and add anything new you find.” For fan-outs, its example prompt adds: “When a subagent reports back, check its evidence before you accept it.” The guide never says what evidence is.
This piece does. By Tuesday your TASKS.md holds one line per item and one receipt per tick: the command that proves it, its exit code, the commit it ran against, the log it wrote, the model that actually served the turn and the subagent that did the work. A verifier re-runs the receipts instead of reading transcripts, and any tick without a receipt goes back to open. The overnight fan-out ends with a table you can check in ten minutes rather than a transcript you have to trust.
The checklist moved out of the harness and into the repo in the last six weeks. Rules for a file in your repo are yours to write.
Aug 14 to Sep 22: the todo list left the harness, and TASKS.md took its place
Claude Code used to keep the checklist for you. Version 2.1.233, on Aug 14, 2026, took it away from the newest models: “Todo/task-tracking tools (TaskCreate/Get/Update/List, TodoWrite) are no longer available on Opus 4.8, Sonnet 5, Fable 5, Mythos 5, and newer models; set CLAUDE_CODE_ENABLE_TODO_TOOLS=1 to bring them back”. Version 2.1.268, on Sep 10, turned that into an allowlist, with the tools “offered only on Claude 3.x, Opus 4.0–4.7, Sonnet 4.0–4.6, Haiku 4.5”.
The tools reference gives the reasoning: “On newer models, Claude keeps track of multi-step work without a written checklist, and the tools’ definitions and reminders take up context.” It also lists exceptions that matter in a mixed fleet. “In background sessions and in cloud sessions, Claude Code provides the same tools on every model, listed or not.”
A model ID Claude Code doesn’t recognize, such as a custom name behind an LLM gateway, gets none. A subagent gets them only when its session has them. To opt back in, use CLAUDE_CODE_ENABLE_TODO_TOOLS=1, --allowedTools TaskCreate, --tools, or the Agent SDK’s allowedTools and tools.
Screenshot: Claude Code Docs, “Tools reference” (undated page), captured Sep 28, 2026.
Twelve days later, Anthropic’s claude.dev guide to Opus 5.5, written by Addy Osmani, put the list somewhere every model can reach. “For a run that will take a while, ask Opus 5.5 to keep its task list in a file and update it as it goes. Then read the file, not the scrollback, to see where the run is.” The reason is compaction: “A long run fills the context window, and Claude Code then summarizes older turns. A list in a file survives that, and it shows you at a glance what’s done and what’s left.”
Screenshot: claude.dev Blog, “Getting the most out of Opus 5.5 in Claude and Claude Code” (Sep 22, 2026), captured Sep 28, 2026.
The same guide treats fan-out as routine. “Early testers had Opus 5.5 coordinate parallel subagents on long audits and migrations, with little oversight.” Its example audit prompt is the template for everything below: “Audit every service in services/ for the retry bug in the linked issue. Give each service to its own subagent. When a subagent reports back, check its evidence before you accept it. Finish with one table: service, affected yes or no, and the evidence.”
Two things the guide doesn’t do. It doesn’t define evidence. And TASKS.md isn’t a Claude Code feature; no Claude Code doc mentions it. It is a file name inside an example prompt, which is what makes it portable: every CLI in your fleet can read and write a Markdown file.
A team with a week of testing behind it landed near the same place from the other side: define done before the run starts.
A tick is a claim the agent makes about its own work
Chatbots suggest; agents act, and a tick is an action. In a fan-out the coordinator reads TASKS.md to decide what’s left, so a false tick is work that never gets scheduled again.
A file survives summarization, which is the point of it. It also survives being wrong. The file keeps whatever the agent wrote, and nothing about a Markdown checkbox knows whether the claim behind it is true.
A subagent’s report has the same problem one level up. It is a summary of the worker, written by the worker, read by a coordinator that can see the work only through that summary. This month’s safety-fallback reports showed how far a summary can drift: a parent crediting one model for work its child did on another, which the served-model log exists to catch. Evidence, in the sense a fleet needs, is anything a second process can re-check without asking the first one.
What must survive compaction is its own policy question, covered in agent-owned compaction. This piece is about what a surviving tick has to carry.
Step 1: Make TASKS.md the record, and put the rule where every CLI reads it
Decide which list is authoritative before the run, because some sessions will have two. Background and cloud sessions still get the built-in task tools on every model; interactive sessions on newer models don’t get them by default. Treat the built-in list as scratch and TASKS.md as the record.
Four dated changes in under six weeks. The harness list went allowlist-only, and the advice moved the list into a file.
Then write three rules into the run contract, the instruction file every CLI in the fleet loads. The run-contract piece covers making each CLI actually read it.
- TASKS.md at the repo root is the record. No other list counts at merge.
- No tick without a receipt, in the schema below.
- The verifier has the last word. A tick it can’t re-run goes back to open, with the reason on the line.
One gap needs a workaround. The sub-agents docs say Claude Code’s built-in Explore and Plan subagents skip CLAUDE.md and any AGENTS.md loaded as project instructions, and so does any subagent with omitClaudeMd. The docs’ advice is to restate a rule that must reach a subagent “in the prompt you give Claude when delegating.” Put the three rules in every delegation prompt, not only in the file.
Step 2: Define the receipt as six fields a script can check
A receipt is the smallest record that lets a second process reproduce the claim. Each field points at something outside the model that wrote it:
| Field | Example | Verified how |
|---|---|---|
cmd |
pnpm --filter billing test retry |
The verifier re-runs it in a clean checkout |
exit |
0 |
Compared with the re-run’s exit code |
sha |
4e1c9a2 |
git cat-file -e finds the commit, and the re-run checks it out |
log |
.runs/billing/retry.log |
Exists, isn’t empty, and its first line names the same sha |
model |
claude-opus-5-5 |
Matches the served-model log for that subagent’s turns, never the worker’s self-report |
agent |
sa-billing-01 |
Matches an ID in the harness’s own records, such as a SubagentStop entry |
Three design choices keep the schema honest. Every field points at a process, a commit, a file or a record. A field the verifier can’t check doesn’t belong in the receipt, however informative it is. And the test wrapper, not the agent, writes the SHA into the log’s first line, so yesterday’s log can’t pass for today’s.
model and agent are the fields teams skip, and they answer the audit questions. model comes from the served-model log, because a subagent that fell back mid-run can still describe itself as the model it was asked to be. agent comes from the harness: in Claude Code, a SubagentStop hook receives each subagent’s agent_transcript_path, per the hooks reference, so the harness itself records which subagents ran and where their transcripts live. A receipt naming an agent the harness never started is forged or confused, and either way the tick reverts.
Items a worker couldn’t confirm get no receipt and no tick. The claude.dev guide’s own advice is “Ask it to mark what it couldn’t confirm”, so give that state its own box, [?], and an honest gap never looks like a finished item.
Step 3: Write TASKS.md so a script can parse it
The format has one job: let the verifier find every tick and its receipt without guessing. Use one line per item, a receipt on the line under each tick, key=value pairs, and three boxes: [ ] open, [x] done with a receipt, [?] couldn’t confirm. A tick the verifier sends back becomes [ ] again, with the reason on the line. The file below is illustrative, built on the claude.dev guide’s retry-bug audit.
# TASKS.md: retry-bug audit, issue 412
Rule: no tick without a receipt. The verifier reverts any tick it cannot re-run.
- [x] services/billing: affected, fixed
receipt: cmd="pnpm --filter billing test retry" exit=0 sha=4e1c9a2 log=.runs/billing/retry.log model=claude-opus-5-5 agent=sa-billing-01
- [x] services/auth: not affected, retry path covered by a test
receipt: cmd="pnpm --filter auth test retry" exit=0 sha=9b07d3e log=.runs/auth/retry.log model=claude-opus-5-5 agent=sa-auth-02
- [?] services/notify: could not confirm, no test target reaches the retry path (agent=sa-notify-04)
- [ ] services/search: affected, fix in progress (agent=sa-search-03)
- [ ] services/export: reverted by verifier at 02:14Z, log missing (ticked by sa-export-05)
- [ ] NEW services/billing: invoice retry spec flaky under load (found by sa-billing-01)
Two details carry weight. New findings arrive as open items marked NEW with the finder’s ID, which turns the guide’s “add anything new you find” into something you can audit. And the reversal note stays on the line, so the next pass, and the person reading in the morning, can see why a tick went back.
Step 4: Fan out one subagent per service, then merge receipts into one table
The guide’s audit prompt already holds the fan-out rule and the finish line. Add the parts that make receipts mergeable. The general fleet patterns are in subagent orchestration; these are the evidence rules that sit on top of them.
| Rule | Why | What breaks without it |
|---|---|---|
| One subagent per service | The guide’s prompt gives each service to its own subagent | Two workers edit one service, and their receipts collide |
| A worker writes only its own lines | The coordinator merges; nobody ticks another worker’s item | A worker ticks a neighbor’s item from a guess |
New findings append as NEW open items |
“add anything new you find”, with a finder ID attached | Discoveries vanish into a report nobody rereads |
Can’t confirm means [?], never [x] |
The guide asks the model to mark what it couldn’t confirm | Guesses turn into ticks |
| Restate the rules in each delegation prompt | Explore and Plan subagents skip the instruction files | Built-in subagents never see the rule |
| Receipts merge into one table at the end | The guide’s finish: service, affected yes or no, and the evidence | The final report turns back into prose |
| No receipt, no tick | Enforced by the verifier in step 5 | The coordinator skips work nobody proved |
The merged table is the guide’s finish line with one column added. Service, affected and evidence become service, affected, receipt and verifier result. The coordinator builds it from TASKS.md, not from the workers’ messages, so the table can’t claim more than the file proves. Workers still report back in prose, which is fine for reasoning and useless as proof.
Step 5: Verify by re-running receipts, not by reading transcripts
The worker proposes the tick. The re-run decides whether it stays.
The verifier is a plain script with no model in the loop. It reads TASKS.md, checks the cheap fields of every [x] first and re-runs the command last. The sketch below is illustrative; adapt the parsing to your file and the re-run to your sandbox.
# verify_tasks.py (illustrative): re-run every receipt, revert ticks that fail
import pathlib, re, shlex, subprocess
TICK = re.compile(r"^- \[x\] (.+)\n receipt: (.+)$", re.M)
AGENTS = set(pathlib.Path(".runs/agents.txt").read_text().split()) # IDs from SubagentStop records
SERVED = dict(l.split() for l in open(".runs/served-models.txt")) # agent -> served model
def check(r):
if subprocess.run(["git", "cat-file", "-e", r["sha"] + "^{commit}"]).returncode:
return "commit not found"
log = pathlib.Path(r["log"])
first = log.read_text().splitlines()[:1] if log.is_file() else []
if not first or r["sha"] not in first[0]:
return "log missing or from another commit"
if r["agent"] not in AGENTS:
return "agent unknown to the harness"
if SERVED.get(r["agent"]) != r["model"]:
return "served model differs"
wt = f"/tmp/verify-{r['sha']}"
subprocess.run(["git", "worktree", "add", "--detach", wt, r["sha"]], check=True)
try:
rerun = subprocess.run(shlex.split(r["cmd"]), cwd=wt, timeout=900)
except subprocess.TimeoutExpired:
return "re-run timed out"
finally:
subprocess.run(["git", "worktree", "remove", "--force", wt])
return None if str(rerun.returncode) == r["exit"] else "re-run exit differs"
text = pathlib.Path("TASKS.md").read_text()
for m in TICK.finditer(text):
r = dict(kv.split("=", 1) for kv in shlex.split(m.group(2)))
reason = check(r)
if reason:
text = text.replace(m.group(0), f"- [ ] {m.group(1)}: reverted by verifier, {reason}")
print("REVERT", m.group(1), "|", reason)
pathlib.Path("TASKS.md").write_text(text)
Order matters. The cheap checks catch the common bad receipts in milliseconds: a commit that doesn’t exist, a log from another run, an agent nobody started. The re-run is the expensive check, and the only one a confident worker can’t talk its way past. On lanes that accept logged fallbacks, downgrade the served-model check from a revert to a flag for re-review.
Code review is moving the same way. GitHub’s Sep 11 changelog says Copilot code review now has “more ways to validate the code under review (e.g., running build commands, running tests, executing targeted scripts, and retrieving information from available tools and APIs).” Read it as an analogy for a reviewer that executes instead of reading; the post says nothing about receipts. Once a reviewer executes, where it runs and what it can see become the questions, and the review-bot piece answers them.
The verifier is a gate, and gates fail. It executes commands an agent wrote, so run it in the same sandbox as the workers: no secrets, no deploy credentials, network only where tests need it. If it crashes or times out, the ticks it didn’t reach count as unverified, and unverified means open at merge. Behind it sits the wall that doesn’t depend on any script: branch protection and a human approval on the pull request, which the PR review policy keeps in human hands.
Step 6: Revert unreceipted ticks on every pass, and finish at zero reverts
Run the verifier at three points: when each subagent stops, before the coordinator builds the final table, and in CI on the pull request. In Claude Code, a SubagentStop command hook is the natural trigger for the first. Reverting is automatic and re-ticking isn’t. A worker that wants an item back at [x] produces a new receipt, and the verifier checks that one too.
The run’s finish line follows: the verifier exits 0 with zero reverted ticks, and every item is [x], or [?] with a named owner. That condition can be proved in a transcript, which is what a goal’s finish line needs when the judge reads nothing else. Make the verifier’s summary the last thing the run prints.
Then count reverts per lane each week. A lane whose ticks revert often has a prompt problem or a model problem, and the model column tells you which to check first.
Where TASKS.md receipts fail, and the signal for each
The stale log. The path exists, and the contents come from an earlier run. Signal: the log’s first line names a different SHA. Fix: the test wrapper writes the SHA and the verifier refuses a mismatch, which it already does if you kept step 2’s rule.
The command that proves nothing. A test filter that matches zero tests exits 0, and so does true. Signal: re-runs that pass in under a second, or logs that report zero tests. Fix: allowlist the commands a receipt may name for each kind of item, and make the wrapper fail on an empty test run.
The tick from the wrong worker. A receipt’s agent isn’t the item’s owner. Signal: the receipt’s ID differs from the one on the item line. Fix: revert it. One subagent per service means one writer per line.
Two lists, one run. A background or cloud session has the built-in task tools on every model, so the agent tracks progress there while TASKS.md goes stale. Signal: the file hasn’t changed in an hour while the session reports progress. Fix: the run contract names the file as the record, and every delegation prompt repeats it.
The swapped worker. A subagent fell back to another model mid-run, and its receipt still names the one it was asked to be. Signal: the receipt’s model differs from the served-model log. Fix: flag it for re-review, or revert it on strict lanes, as step 5’s check does.
The verifier nobody runs. Ticks pile up with no verifier pass after them. Signal: [x] lines newer than the last verifier summary. Fix: CI runs the verifier on the pull request, and merge treats unverified ticks as open.
Receipts are how a fleet remembers what it proved
A transcript records what an agent said. A receipt records what a second process could reproduce, and across a fleet of CLIs only the second survives a Monday review. Receipts in a file, checked by a script and merged into one table, are evidence an operating layer can hold for every lane at once, which is the job a multi-agent command center exists to do. The models behind your lanes will keep changing; the rule that a tick needs a receipt doesn’t have to.
FAQ
Does Claude Code still have a built-in todo list?
Only on some models and sessions. Since 2.1.268 (Sep 10, 2026), the task tools are offered by default only on Claude 3.x, Opus 4.0–4.7, Sonnet 4.0–4.6 and Haiku 4.5. Background and cloud sessions still get them on every model, and CLAUDE_CODE_ENABLE_TODO_TOOLS=1 brings them back elsewhere.
What counts as evidence when a subagent reports back?
Anthropic’s guide doesn’t define it. A workable rule: anything a script can re-check without trusting the model. That means the command, its exit code, the commit it ran against, the log it wrote, the served model and the subagent’s ID. A verifier re-runs the command; a tick it can’t reproduce goes back to open.
Sources
- claude.dev: Getting the most out of Opus 5.5 in Claude and Claude Code (Addy Osmani, Sep 22, 2026)
- Claude Code: Tools reference, “Task tool availability”
- Claude Code changelog (2.1.233, Aug 14, 2026; 2.1.268, Sep 10, 2026)
- Claude Code: Subagents
- Claude Code: Hooks reference,
SubagentStop - GitHub changelog: Copilot code review analysis updates (Sep 11, 2026)
