Review Bots Now Run PR Code: Where They Execute, What They See, What a Verdict Proves

Review bots now run PR code. Register where each one executes, which secrets and egress it gets and what its verdict proves, then drill it with a canary PR.

Review bots hero: a bot's review card running build and test commands, tagged with three open questions, runner, secrets and egressReview bots hero: a bot's review card running build and test commands, tagged with three open questions, runner, secrets and egress
The review now runs. The question is where, and with what.

One sentence in GitHub’s code-review docs deserves a second read: “When reviewing a pull request, Copilot reads repository custom instructions, agent instructions, and agent skills from the head branch (the branch with your changes), not the base branch.” Since Sep 11, that same reviewer can run build commands, tests and targeted scripts. Put the two together and a pull request can rewrite the instructions of the bot that executes it.

That’s the new job description for review bots. They used to read a diff and leave comments. Now at least one of them can run your code on a runner, with an environment, a network and whatever that environment holds. So the questions moved: where does each bot execute, what can it see, and what does its verdict actually prove?

The move for Tuesday is a register, three rules and a drill. Write one row per bot. Keep self-hosted runners and deploy secrets away from anything that builds code you didn’t write, and leave approval-counting off. Grade every verdict by its receipts.

Then send each bot a canary pull request and let your own logs fill the rows the vendors leave blank. Chatbots suggest; agents act, and a reviewer that acts is a CI job with opinions.

What changed for review bots between April and September

GitHub, Sep 11. Copilot code review “now uses the full set of shell tools from the Copilot SDK, running behind the Copilot agent firewall.” The changelog lists what that buys: “running build commands, running tests, executing targeted scripts, and retrieving information from available tools and APIs”. The same post adds an ensemble, scoped to one level: “The Lite effort level now uses an ensemble of agents to produce a review rather than one agent working alone.”

GitHub changelog section titled Deeper analysis with shell tools, saying Copilot code review now uses the full set of shell tools from the Copilot SDK behind the Copilot agent firewall to run builds, tests and targeted scripts Screenshot: GitHub Changelog, “Auto-resolution and analysis updates in Copilot code review - GitHub Changelog” (Sep 11, 2026), captured Sep 28, 2026.

The post never mentions runners. The code-review docs do: “By default, Copilot code review uses standard GitHub-hosted runners.” Larger hosted runners and self-hosted runners are options, and the runner docs narrow the second one: “ARC is the only officially supported solution for self-hosting Copilot code review. For security reasons, do not use non-ARC self-hosted runners.” The code-review docs add that “The GitHub MCP server and Playwright MCP server are enabled by default.”

The firewall is on by default, with limits GitHub states itself. The firewall docs say it “only applies to processes started by the agent via its Bash tool”, not to MCP server processes or setup-step processes, and “should not be considered a comprehensive security solution.” Its recommended allowlist, on by default, admits OS package repositories, container registries and language package registries. By default, repository admins can add their own allow rules.

Then there’s the head-branch sentence from the top. In July, GitHub’s changelog account said where code review’s skills live, and that location travels with every branch.

GitHub, Sep 1. Copilot’s review can now carry an approval that counts. “When enabled, Copilot can submit an approval that counts toward the repository’s required-approvals rule.” The announcement calls it a public preview and says “The ability for Copilot to approve is off by default and configurable at the enterprise, organization, and repository level.” New commits dismiss the approval, as they would a human’s.

Cursor, Apr 30 and Sep 23. Cursor put Security Review in beta on Teams and Enterprise plans in April. Its Security Reviewer checks each PR for, among other things, “agent tool auto-approvals, and prompt injection attacks”. On Sep 23, Cursor relaunched it as one of two Cursor bots, beside Rollouts. It “posts one review comment reporting exploitable bugs.”

The Security Agents docs say “Both agent types run on the Automations platform and require Cloud Agents”, under a shared team service account. None of these Cursor pages says whether Security Review builds or executes PR code.

Anthropic, Sep 10 and Sep 25. Claude Code’s Code Review docs say “multiple agents analyze the diff and surrounding code in parallel on Anthropic infrastructure.” A verification step then checks candidates against actual code behavior. The verdict never gates a merge: “The check run always completes with a neutral conclusion so it never blocks merging through branch protection rules.”

Forks wait for a person: “Claude doesn’t review a pull request from a fork automatically, regardless of the repository’s Review Behavior setting.” A top-level @claude review comment starts one, from someone with write, maintain or admin permission on the base repository. Re-run and new pushes don’t.

Claude Code Docs page on Code Review, How reviews work section, saying multiple agents analyze the diff and surrounding code in parallel on Anthropic infrastructure and a verification step checks candidates against actual code behavior Screenshot: Claude Code Docs, “Code Review - Claude Code Docs” (undated page), captured Sep 28, 2026.

Two releases changed what a verdict means. The CHANGELOG for v2.1.268 (Sep 10) says a review whose verifying agent fails midway “now replaces that agent and reaches a verdict”. v2.1.283 (Sep 25) fixed billing for “a review that stopped at its time limit with nothing verified: it now shows as incomplete, isn’t charged, and is retried once”. The docs don’t say whether the verification step runs PR code.

Step 1: Write the review-bot register, one row per bot

One row per bot, one column per question, including the questions vendors leave blank. A blank is information: it’s the cell the canary fills in step 4. The rows below come from vendor docs and changelogs read Sep 28. “Not documented” means the docs don’t say, not that the answer is no.

Bot Where it executes What it may run Secrets and env it sees Egress Fork PRs Approval counts? Reads its instructions from
Copilot code review GitHub Actions: standard hosted runners by default; larger hosted, or ARC self-hosted Shell tools: builds, tests, targeted scripts; GitHub and Playwright MCP on by default Not documented Agent firewall on by default, for Bash-tool processes only; allowlist admits package registries Not documented Only if an admin enables it (public preview, off by default) The PR’s head branch
Claude Code Code Review Anthropic infrastructure Parallel agents plus a verification step; running PR code not documented Not documented Not documented Never automatic; @claude review from write, maintain or admin on the base repo No: neutral check run, never blocks Not documented
Cursor Security Review Cloud Agents on the Automations platform, under a shared team service account Needs at least one tool or MCP; executing PR code not documented Not documented Not documented Not stated on the Security Agents page No approval documented; one review comment Not documented
Your own Actions-based bot Your runners What the workflow grants What the workflow grants Your egress policy Your trigger rules Your branch protection The ref your workflow checks out

Two columns need more than a cell.

Reads its instructions from. Copilot answers it: the head branch. A PR that edits the repository’s instructions, or anything under .github/skills, edits the reviewer that grades it, and since Sep 11 that reviewer runs commands. Give instruction and skill files a human code owner, and treat the bot’s verdict on any PR that touches them as advisory until that owner approves the change. Anthropic’s docs don’t say which branch Code Review reads REVIEW.md from, so that cell waits for the drill.

Fork PRs. Claude Code Review has the strictest documented answer: never automatic, and whoever asks needs write access to the base repository. GitHub’s docs, read Sep 28, do not say what Copilot code review does with fork PRs. Cursor’s Automations docs say pull request triggers don’t run on PRs opened from forks, but its Security Agents page doesn’t restate that for Security Review. Until your canary says otherwise, read those blanks as yes.

Who may trigger a workflow at all belongs on the Actions trigger allowlist. This column is about what the bot can reach once something has triggered it.

Diagram of the review-bot trust boundary: the PR head branch supplies code, instructions and skills to the review-bot runtime of runner, env vars, secrets, tools and egress, which produces a verdict that feeds the merge gate, while a canary PR probes the runtime and its results fill the review-bot registerDiagram of the review-bot trust boundary: the PR head branch supplies code, instructions and skills to the review-bot runtime of runner, env vars, secrets, tools and egress, which produces a verdict that feeds the merge gate, while a canary PR probes the runtime and its results fill the review-bot register The author controls everything on the left of the boundary. The register records what sits inside it.

Step 2: Hold review bots to three rules, written as policy

  1. No self-hosted runners for bots that build fork PRs. GitHub already narrows self-hosting for Copilot to ARC. Go one step further: nothing that might build a fork runs on a runner inside your network, and GitHub doesn’t document whether Copilot builds forks. Every runner still gets the control GitHub asks for: “Configure network security controls for your GitHub Actions runners to ensure that Copilot code review does not have open access to your network or the public internet.”
  2. No deploy secrets in review environments. GitHub’s docs do not say which secrets, if any, reach the review’s environment, but they do document sharing: org owners can set one runner type for code review and the cloud agent together, and the two share the repository’s MCP configuration. So assume anything the cloud agent can read, the reviewer can read too. Nothing there should be able to deploy, publish a package or change cloud resources. Check MCP servers separately, because the firewall doesn’t cover their processes.
  3. Approval-counting off. Leave Copilot’s approval switch at its default, off. The PR review agent policy already rules out a bot approving a bot’s work; Sep 1 is the day a vendor shipped the switch that tests it. If a repository ever needs an exception, the docs let it limit counted approvals to file globs, up to 15. Give that exception an owner and an expiry date.

With the head-branch rule and step 3’s verdict rule added, the policy fits in one file. The shape is illustrative and the field names are ours.

# review-bots.policy.yaml (illustrative; field names are ours)
review_bots:
  runners:
    fork_prs_on_self_hosted: never
    copilot_self_hosted: arc_only          # GitHub: no non-ARC runners
    runner_network: allowlist_only         # no open path to your network or the internet
  secrets:
    in_review_environment: no_deploy_power # deploy, publish, cloud-admin: never
    cloud_agent_secrets: treat_as_reviewer_visible
    mcp_server_secrets: review_separately  # the firewall does not cover MCP processes
  approvals:
    bot_approval_counts: false
    exceptions: {owner: required, expires: required, max_globs: 15}
  instructions:
    record_config_source: [head, base, unknown]
    instruction_and_skill_changes: human_code_owner_required
  verdicts:
    verified_only_with: [command_and_exit_code, named_tests_on_head_commit, reproduced_finding]
    incomplete_means: no_review
  drill:
    canary_pr_from: [branch, fork]
    rerun_after: any vendor change to the bot

Rules like these fail quietly. By default a repository admin can add a firewall allow rule, and an org owner can switch the runner type for code review and the cloud agent in one move. Neither change pages anyone. The canary in step 4 is how you notice.

The wall behind it all is rule 2: a runner with nothing worth stealing, on a network that goes nowhere, caps the damage when another rule slips. Firewalls and sandboxes are guardrails; the Black Hat sandbox escapes were the reminder.

Step 3: Grade each verdict by its receipts

A verdict is a claim. What turns it into evidence is what turns a ticked task into evidence: a receipt you can re-run. TASKS.md with receipts sets that standard for an agent’s own work, and review bots get no discount.

Grade every verdict against this table. The quoted lines are the vendors’; the grades are ours.

What the verdict shows Counts as verified? Why
A command, its exit code and its log Yes, if you can re-run it on the same commit The unit of evidence
Named tests with a result for each, on the head commit Yes Check the commit, not the branch name
A reproduced finding: a failing test, a script’s output, a request and its response Yes, the strongest receipt Anyone can re-check it
Line comments with no commands shown No; it’s a read Useful, but review by reading
A Copilot approval assessment on its own No GitHub: “An approval assessment alone does not count toward merge requirements.”
Claude Code Review’s neutral check run No, by design It “never blocks merging”
Incomplete, nothing verified No review happened Not a pass; since v2.1.283 it’s retried once
A verdict on a PR that changed the reviewer’s own instructions Advisory only Until a human code owner approves the change

Two notes on using it. GitHub doesn’t say which commands its review agent chose to run on a given PR, or whether it shows you that list. If a bot can’t show its commands, grade its verdict as a read, however confident it sounds. And don’t average grades: one reproduced finding outweighs forty unverified comments, and a dashboard that counts reviews without grading them will count the incomplete ones as green.

Step 4: Send every bot a canary PR, from a branch and from a fork

The drill turns “not documented” into “observed”. The script below is illustrative: wire it in as the test command on the canary PR only, so any bot that runs the tests runs it. It prints environment variable names, never values, and it probes egress. Point the second URL at a host you control; its access log is a receipt the bot can’t edit.

#!/usr/bin/env bash
# canary.sh (illustrative): the test command on the canary PR only.
# Prints environment variable NAMES, never values, then probes egress.
set -u
echo "canary $(date -u +%Y-%m-%dT%H:%M:%SZ) host=$(uname -n)"

echo "== env var names (values are never printed)"
compgen -e | sort

echo "== names that look like credentials"
compgen -e | grep -Ei 'token|secret|key|pass|cred|deploy' || echo "none"

echo "== egress probes (000 = no connection)"
for url in https://registry.npmjs.org/ \
           https://canary.example.com/ping \
           https://build.internal.example.com/; do
  code=$(curl -s -o /dev/null -m 5 -w '%{http_code}' "$url" || true)
  echo "$url -> ${code:-000}"
done

It uses compgen -e because that lists exported names and nothing else. Parsing env output instead can leak part of a multi-line value.

Run the drill the same way for every bot:

  1. Open the canary twice. Once from a branch in the repository, once from a fork. Request a review the way your team normally does. To fill the instructions column, let the canary also edit the bot’s instruction file, such as REVIEW.md for Claude Code Review, to ask for a harmless marker word; a verdict that uses the word came from the PR’s copy.
  2. Record what ran. Did the bot run the tests? Did its verdict mention the canary’s output? Did your canary host log a request, and from which address?
  3. Fill the register. Each “not documented” cell becomes “observed” or “no evidence”, with the date. No evidence isn’t the same as no. A bot that never ran the canary proved nothing about what it could see.
  4. Fix what surprised you against step 2’s rules: a credential-shaped name, a host you didn’t expect to reach, a fork that got a full review.
  5. Re-run after every vendor change to the bot. Those changes arrive through the intake rule in the admin-toggle sweep, which already lists the Copilot Code Review policy as a toggle to decide.

The cadence matters because the bots keep moving. Seven dated changes in five months is the case for step 5.

Timeline chart of review bot changes by vendor from Apr 30 to Sep 25, 2026: Cursor Security Review beta Apr 30 and relaunch Sep 23, Copilot code review skills and MCP Jul 29, approvals Sep 1 and shell tools Sep 11, Claude Code Review fixes Sep 10 and Sep 25Timeline chart of review bot changes by vendor from Apr 30 to Sep 25, 2026: Cursor Security Review beta Apr 30 and relaunch Sep 23, Copilot code review skills and MCP Jul 29, approvals Sep 1 and shell tools Sep 11, Claude Code Review fixes Sep 10 and Sep 25 Dates from GitHub’s and Cursor’s changelogs, the @GHchangelog post, and npm publish times for Claude Code. Read Sep 28, 2026.

Where review-bot controls fail, and the signal for each

  • A PR edits the reviewer’s instructions. Signal: instruction files or anything under .github/skills in the diff. First move: hold the bot’s verdict as advisory until the code owner signs off.
  • A firewall allow rule appears. Signal: a host that returned 000 in last month’s canary now connects. First move: find who added the rule, and why, before the next fork PR lands.
  • The runner type changes. Signal: the canary’s host line changes, or a self-hosted runner label shows up in the run. First move: confirm it’s ARC and that it sits outside your network.
  • An MCP server holds a credential. Signal: MCP configuration that references a secret. The firewall doesn’t cover MCP processes, and the canary can’t see their environment. First move: move the secret out, or keep that server out of review.
  • Incomplete reviews get counted as passes. Signal: green review totals with no receipts behind them. First move: re-grade last month’s verdicts with step 3’s table.
  • A bot approval starts counting. Signal: a merge that met its required approvals with a bot’s approval among them. First move: switch it off, then review everything it let through.

Every one of these shows up in your own evidence, the diff, the canary log or the merge record, and in none of the vendors’ changelogs. That’s the reason to keep the evidence yourself.

A review bot is a fleet lane that other people’s code can start

Fleet operators ask four questions of every agent lane: who owns it, what it can touch, what it costs and how it stops. A review bot is a lane that a stranger’s pull request can start, which makes the second question the one that matters, and the register is where you answer it. The verdict grades answer a different one: can you reconstruct what the bot did? You can’t replay what you can’t see, and a verdict with no commands behind it is exactly that.

A bot that builds fork PRs is also an intake path for untrusted input, next to error trackers and alert feeds that hand work to agents. The auto-remediation register lists those paths. Give every review bot that builds forks a row there as well.

FAQ

Does Copilot code review run code from fork pull requests?

GitHub’s docs don’t say. None of its code-review, runner, firewall or secrets pages that we checked on Sep 28 mentions forks. They do say reviews run on GitHub Actions, with shell tools behind a firewall. Until GitHub documents it, send a canary pull request from a fork and record what actually runs.

Can a pull request change what Copilot code review checks?

Yes. GitHub’s docs say Copilot reads repository custom instructions, agent instructions and agent skills from the pull request’s head branch, not the base branch. A PR can therefore edit the instructions its own reviewer follows. Give instruction and skill files a human code owner, and treat verdicts on PRs that change them as advisory.

Does Claude Code Review block a merge?

No. Anthropic’s docs say the check run always completes with a neutral conclusion, so it never blocks merging through branch protection rules, and findings don’t approve or block the PR. Since v2.1.283, a review that hit its time limit with nothing verified shows as incomplete and is retried once.

Sources

YOU'RE THROUGH THIS ONE.

Keep connecting the dots.

Back to the library