Review Bots Now Run PR Code: Where They Execute, What They See, What a Verdict Proves
Review bots now run PR code. Register where each one executes, which secrets and egress it gets and what its verdict proves, then drill it with a canary PR.
Go deeper. Build your own.
One sentence in GitHub’s code-review docs deserves a second read: “When reviewing a pull request, Copilot reads repository custom instructions, agent instructions, and agent skills from the head branch (the branch with your changes), not the base branch.” Since Sep 11, that same reviewer can run build commands, tests and targeted scripts. Put the two together and a pull request can rewrite the instructions of the bot that executes it.
That’s the new job description for review bots. They used to read a diff and leave comments. Now at least one of them can run your code on a runner, with an environment, a network and whatever that environment holds. So the questions moved: where does each bot execute, what can it see, and what does its verdict actually prove?
The move for Tuesday is a register, three rules and a drill. Write one row per bot. Keep self-hosted runners and deploy secrets away from anything that builds code you didn’t write, and leave approval-counting off. Grade every verdict by its receipts.
Then send each bot a canary pull request and let your own logs fill the rows the vendors leave blank. Chatbots suggest; agents act, and a reviewer that acts is a CI job with opinions.
What changed for review bots between April and September
GitHub, Sep 11. Copilot code review “now uses the full set of shell tools from the Copilot SDK, running behind the Copilot agent firewall.” The changelog lists what that buys: “running build commands, running tests, executing targeted scripts, and retrieving information from available tools and APIs”. The same post adds an ensemble, scoped to one level: “The Lite effort level now uses an ensemble of agents to produce a review rather than one agent working alone.”
Screenshot: GitHub Changelog, “Auto-resolution and analysis updates in Copilot code review - GitHub Changelog” (Sep 11, 2026), captured Sep 28, 2026.
The post never mentions runners. The code-review docs do: “By default, Copilot code review uses standard GitHub-hosted runners.” Larger hosted runners and self-hosted runners are options, and the runner docs narrow the second one: “ARC is the only officially supported solution for self-hosting Copilot code review. For security reasons, do not use non-ARC self-hosted runners.” The code-review docs add that “The GitHub MCP server and Playwright MCP server are enabled by default.”
The firewall is on by default, with limits GitHub states itself. The firewall docs say it “only applies to processes started by the agent via its Bash tool”, not to MCP server processes or setup-step processes, and “should not be considered a comprehensive security solution.” Its recommended allowlist, on by default, admits OS package repositories, container registries and language package registries. By default, repository admins can add their own allow rules.
Then there’s the head-branch sentence from the top. In July, GitHub’s changelog account said where code review’s skills live, and that location travels with every branch.
GitHub, Sep 1. Copilot’s review can now carry an approval that counts. “When enabled, Copilot can submit an approval that counts toward the repository’s required-approvals rule.” The announcement calls it a public preview and says “The ability for Copilot to approve is off by default and configurable at the enterprise, organization, and repository level.” New commits dismiss the approval, as they would a human’s.
Cursor, Apr 30 and Sep 23. Cursor put Security Review in beta on Teams and Enterprise plans in April. Its Security Reviewer checks each PR for, among other things, “agent tool auto-approvals, and prompt injection attacks”. On Sep 23, Cursor relaunched it as one of two Cursor bots, beside Rollouts. It “posts one review comment reporting exploitable bugs.”
The Security Agents docs say “Both agent types run on the Automations platform and require Cloud Agents”, under a shared team service account. None of these Cursor pages says whether Security Review builds or executes PR code.
Anthropic, Sep 10 and Sep 25. Claude Code’s Code Review docs say “multiple agents analyze the diff and surrounding code in parallel on Anthropic infrastructure.” A verification step then checks candidates against actual code behavior. The verdict never gates a merge: “The check run always completes with a neutral conclusion so it never blocks merging through branch protection rules.”
Forks wait for a person: “Claude doesn’t review a pull request from a fork automatically, regardless of the repository’s Review Behavior setting.” A top-level @claude review comment starts one, from someone with write, maintain or admin permission on the base repository. Re-run and new pushes don’t.
Screenshot: Claude Code Docs, “Code Review - Claude Code Docs” (undated page), captured Sep 28, 2026.
Two releases changed what a verdict means. The CHANGELOG for v2.1.268 (Sep 10) says a review whose verifying agent fails midway “now replaces that agent and reaches a verdict”. v2.1.283 (Sep 25) fixed billing for “a review that stopped at its time limit with nothing verified: it now shows as incomplete, isn’t charged, and is retried once”. The docs don’t say whether the verification step runs PR code.
Step 1: Write the review-bot register, one row per bot
One row per bot, one column per question, including the questions vendors leave blank. A blank is information: it’s the cell the canary fills in step 4. The rows below come from vendor docs and changelogs read Sep 28. “Not documented” means the docs don’t say, not that the answer is no.
| Bot | Where it executes | What it may run | Secrets and env it sees | Egress | Fork PRs | Approval counts? | Reads its instructions from |
|---|---|---|---|---|---|---|---|
| Copilot code review | GitHub Actions: standard hosted runners by default; larger hosted, or ARC self-hosted | Shell tools: builds, tests, targeted scripts; GitHub and Playwright MCP on by default | Not documented | Agent firewall on by default, for Bash-tool processes only; allowlist admits package registries | Not documented | Only if an admin enables it (public preview, off by default) | The PR’s head branch |
| Claude Code Code Review | Anthropic infrastructure | Parallel agents plus a verification step; running PR code not documented | Not documented | Not documented | Never automatic; @claude review from write, maintain or admin on the base repo |
No: neutral check run, never blocks | Not documented |
| Cursor Security Review | Cloud Agents on the Automations platform, under a shared team service account | Needs at least one tool or MCP; executing PR code not documented | Not documented | Not documented | Not stated on the Security Agents page | No approval documented; one review comment | Not documented |
| Your own Actions-based bot | Your runners | What the workflow grants | What the workflow grants | Your egress policy | Your trigger rules | Your branch protection | The ref your workflow checks out |
Two columns need more than a cell.
Reads its instructions from. Copilot answers it: the head branch. A PR that edits the repository’s instructions, or anything under .github/skills, edits the reviewer that grades it, and since Sep 11 that reviewer runs commands. Give instruction and skill files a human code owner, and treat the bot’s verdict on any PR that touches them as advisory until that owner approves the change. Anthropic’s docs don’t say which branch Code Review reads REVIEW.md from, so that cell waits for the drill.
Fork PRs. Claude Code Review has the strictest documented answer: never automatic, and whoever asks needs write access to the base repository. GitHub’s docs, read Sep 28, do not say what Copilot code review does with fork PRs. Cursor’s Automations docs say pull request triggers don’t run on PRs opened from forks, but its Security Agents page doesn’t restate that for Security Review. Until your canary says otherwise, read those blanks as yes.
Who may trigger a workflow at all belongs on the Actions trigger allowlist. This column is about what the bot can reach once something has triggered it.
The author controls everything on the left of the boundary. The register records what sits inside it.
Step 2: Hold review bots to three rules, written as policy
- No self-hosted runners for bots that build fork PRs. GitHub already narrows self-hosting for Copilot to ARC. Go one step further: nothing that might build a fork runs on a runner inside your network, and GitHub doesn’t document whether Copilot builds forks. Every runner still gets the control GitHub asks for: “Configure network security controls for your GitHub Actions runners to ensure that Copilot code review does not have open access to your network or the public internet.”
- No deploy secrets in review environments. GitHub’s docs do not say which secrets, if any, reach the review’s environment, but they do document sharing: org owners can set one runner type for code review and the cloud agent together, and the two share the repository’s MCP configuration. So assume anything the cloud agent can read, the reviewer can read too. Nothing there should be able to deploy, publish a package or change cloud resources. Check MCP servers separately, because the firewall doesn’t cover their processes.
- Approval-counting off. Leave Copilot’s approval switch at its default, off. The PR review agent policy already rules out a bot approving a bot’s work; Sep 1 is the day a vendor shipped the switch that tests it. If a repository ever needs an exception, the docs let it limit counted approvals to file globs, up to 15. Give that exception an owner and an expiry date.
With the head-branch rule and step 3’s verdict rule added, the policy fits in one file. The shape is illustrative and the field names are ours.
# review-bots.policy.yaml (illustrative; field names are ours)
review_bots:
runners:
fork_prs_on_self_hosted: never
copilot_self_hosted: arc_only # GitHub: no non-ARC runners
runner_network: allowlist_only # no open path to your network or the internet
secrets:
in_review_environment: no_deploy_power # deploy, publish, cloud-admin: never
cloud_agent_secrets: treat_as_reviewer_visible
mcp_server_secrets: review_separately # the firewall does not cover MCP processes
approvals:
bot_approval_counts: false
exceptions: {owner: required, expires: required, max_globs: 15}
instructions:
record_config_source: [head, base, unknown]
instruction_and_skill_changes: human_code_owner_required
verdicts:
verified_only_with: [command_and_exit_code, named_tests_on_head_commit, reproduced_finding]
incomplete_means: no_review
drill:
canary_pr_from: [branch, fork]
rerun_after: any vendor change to the bot
Rules like these fail quietly. By default a repository admin can add a firewall allow rule, and an org owner can switch the runner type for code review and the cloud agent in one move. Neither change pages anyone. The canary in step 4 is how you notice.
The wall behind it all is rule 2: a runner with nothing worth stealing, on a network that goes nowhere, caps the damage when another rule slips. Firewalls and sandboxes are guardrails; the Black Hat sandbox escapes were the reminder.
Step 3: Grade each verdict by its receipts
A verdict is a claim. What turns it into evidence is what turns a ticked task into evidence: a receipt you can re-run. TASKS.md with receipts sets that standard for an agent’s own work, and review bots get no discount.
Grade every verdict against this table. The quoted lines are the vendors’; the grades are ours.
| What the verdict shows | Counts as verified? | Why |
|---|---|---|
| A command, its exit code and its log | Yes, if you can re-run it on the same commit | The unit of evidence |
| Named tests with a result for each, on the head commit | Yes | Check the commit, not the branch name |
| A reproduced finding: a failing test, a script’s output, a request and its response | Yes, the strongest receipt | Anyone can re-check it |
| Line comments with no commands shown | No; it’s a read | Useful, but review by reading |
| A Copilot approval assessment on its own | No | GitHub: “An approval assessment alone does not count toward merge requirements.” |
| Claude Code Review’s neutral check run | No, by design | It “never blocks merging” |
| Incomplete, nothing verified | No review happened | Not a pass; since v2.1.283 it’s retried once |
| A verdict on a PR that changed the reviewer’s own instructions | Advisory only | Until a human code owner approves the change |
Two notes on using it. GitHub doesn’t say which commands its review agent chose to run on a given PR, or whether it shows you that list. If a bot can’t show its commands, grade its verdict as a read, however confident it sounds. And don’t average grades: one reproduced finding outweighs forty unverified comments, and a dashboard that counts reviews without grading them will count the incomplete ones as green.
Step 4: Send every bot a canary PR, from a branch and from a fork
The drill turns “not documented” into “observed”. The script below is illustrative: wire it in as the test command on the canary PR only, so any bot that runs the tests runs it. It prints environment variable names, never values, and it probes egress. Point the second URL at a host you control; its access log is a receipt the bot can’t edit.
#!/usr/bin/env bash
# canary.sh (illustrative): the test command on the canary PR only.
# Prints environment variable NAMES, never values, then probes egress.
set -u
echo "canary $(date -u +%Y-%m-%dT%H:%M:%SZ) host=$(uname -n)"
echo "== env var names (values are never printed)"
compgen -e | sort
echo "== names that look like credentials"
compgen -e | grep -Ei 'token|secret|key|pass|cred|deploy' || echo "none"
echo "== egress probes (000 = no connection)"
for url in https://registry.npmjs.org/ \
https://canary.example.com/ping \
https://build.internal.example.com/; do
code=$(curl -s -o /dev/null -m 5 -w '%{http_code}' "$url" || true)
echo "$url -> ${code:-000}"
done
It uses compgen -e because that lists exported names and nothing else. Parsing env output instead can leak part of a multi-line value.
Run the drill the same way for every bot:
- Open the canary twice. Once from a branch in the repository, once from a fork. Request a review the way your team normally does. To fill the instructions column, let the canary also edit the bot’s instruction file, such as
REVIEW.mdfor Claude Code Review, to ask for a harmless marker word; a verdict that uses the word came from the PR’s copy. - Record what ran. Did the bot run the tests? Did its verdict mention the canary’s output? Did your canary host log a request, and from which address?
- Fill the register. Each “not documented” cell becomes “observed” or “no evidence”, with the date. No evidence isn’t the same as no. A bot that never ran the canary proved nothing about what it could see.
- Fix what surprised you against step 2’s rules: a credential-shaped name, a host you didn’t expect to reach, a fork that got a full review.
- Re-run after every vendor change to the bot. Those changes arrive through the intake rule in the admin-toggle sweep, which already lists the Copilot Code Review policy as a toggle to decide.
The cadence matters because the bots keep moving. Seven dated changes in five months is the case for step 5.
Dates from GitHub’s and Cursor’s changelogs, the @GHchangelog post, and npm publish times for Claude Code. Read Sep 28, 2026.
Where review-bot controls fail, and the signal for each
- A PR edits the reviewer’s instructions. Signal: instruction files or anything under
.github/skillsin the diff. First move: hold the bot’s verdict as advisory until the code owner signs off. - A firewall allow rule appears. Signal: a host that returned 000 in last month’s canary now connects. First move: find who added the rule, and why, before the next fork PR lands.
- The runner type changes. Signal: the canary’s host line changes, or a self-hosted runner label shows up in the run. First move: confirm it’s ARC and that it sits outside your network.
- An MCP server holds a credential. Signal: MCP configuration that references a secret. The firewall doesn’t cover MCP processes, and the canary can’t see their environment. First move: move the secret out, or keep that server out of review.
- Incomplete reviews get counted as passes. Signal: green review totals with no receipts behind them. First move: re-grade last month’s verdicts with step 3’s table.
- A bot approval starts counting. Signal: a merge that met its required approvals with a bot’s approval among them. First move: switch it off, then review everything it let through.
Every one of these shows up in your own evidence, the diff, the canary log or the merge record, and in none of the vendors’ changelogs. That’s the reason to keep the evidence yourself.
A review bot is a fleet lane that other people’s code can start
Fleet operators ask four questions of every agent lane: who owns it, what it can touch, what it costs and how it stops. A review bot is a lane that a stranger’s pull request can start, which makes the second question the one that matters, and the register is where you answer it. The verdict grades answer a different one: can you reconstruct what the bot did? You can’t replay what you can’t see, and a verdict with no commands behind it is exactly that.
A bot that builds fork PRs is also an intake path for untrusted input, next to error trackers and alert feeds that hand work to agents. The auto-remediation register lists those paths. Give every review bot that builds forks a row there as well.
FAQ
Does Copilot code review run code from fork pull requests?
GitHub’s docs don’t say. None of its code-review, runner, firewall or secrets pages that we checked on Sep 28 mentions forks. They do say reviews run on GitHub Actions, with shell tools behind a firewall. Until GitHub documents it, send a canary pull request from a fork and record what actually runs.
Can a pull request change what Copilot code review checks?
Yes. GitHub’s docs say Copilot reads repository custom instructions, agent instructions and agent skills from the pull request’s head branch, not the base branch. A PR can therefore edit the instructions its own reviewer follows. Give instruction and skill files a human code owner, and treat verdicts on PRs that change them as advisory.
Does Claude Code Review block a merge?
No. Anthropic’s docs say the check run always completes with a neutral conclusion, so it never blocks merging through branch protection rules, and findings don’t approve or block the PR. Since v2.1.283, a review that hit its time limit with nothing verified shows as incomplete and is retried once.
Sources
- GitHub changelog, Sep 11, 2026: auto-resolution and analysis updates in Copilot code review — shell tools behind the agent firewall; the Lite ensemble
- GitHub changelog, Sep 1, 2026: Copilot code review can now approve pull requests — approvals that count, off by default, public preview
- GitHub Docs: Copilot code review — Actions runners, head-branch instructions and skills, MCP defaults
- GitHub Docs: configure runners — ARC-only self-hosting and runner network controls
- GitHub Docs: customize the firewall — defaults, the recommended allowlist and stated limitations
- Cursor changelog, Apr 30, 2026 — Security Review in beta on Teams and Enterprise
- Cursor changelog, Sep 23, 2026: Rollouts and Security Review — the relaunch as one of two Cursor bots
- Cursor docs: Security Agents — Cloud Agents on the Automations platform, shared team service account
- Cursor docs: Automations — pull request triggers don’t run on fork PRs
- Claude Code docs: Code Review — Anthropic infrastructure, neutral check run, fork PRs
- Claude Code CHANGELOG — the v2.1.268 and v2.1.283 Code Review fixes
