Add a Critic, Not a Second Reviewer: Risk-Route AI Code Review Agents Like Meta's RADAR

Two AI reviewers scored worst in a Cornell/Stanford test; a critic of the review scored best. Route each PR by risk, as Meta's RADAR does, with this protocol.

Hero illustration for AI code review agents: one pull request card splits into four lanes of increasing risk, and the two highest-risk lanes gain a critic node that inspects the reviewer rather than the code, and the top lane ends at a human gateHero illustration for AI code review agents: one pull request card splits into four lanes of increasing risk, and the two highest-risk lanes gain a critic node that inspects the reviewer rather than the code, and the top lane ends at a human gate
Route first, then decide who reads. The critic sits on the review, not on the diff.

Seventy-five percent. That was the score for two independent AI reviewers checking the same code in a Cornell/Stanford experiment published in August, and it was the lowest result of every setup the authors tried. One reviewer alone did better. A reviewer plus a third agent whose only job was to audit the review did best of all, at 87%.

If your plan for the flood of agent-written pull requests is to bolt a second AI reviewer onto the first, that result is the one to sit with. More reviewers is the instinct. A better shape is two moves: decide how much review a PR deserves before anyone reads it, and where one AI reviewer is not enough, add a critic that grades the review instead of a second opinion on the code.

This piece turns that into a risk-routed review protocol for AI code review agents: four risk tiers, who reviews each, what the critic may do, when a disagreement goes to a human, what each PR may spend, and what evidence it leaves behind. Meta’s RADAR supplies the routing half. The Adversarial Review paper supplies the critic. They are different papers, and the protocol only works if you keep them apart.

Meta RADAR and the critic paper: two results that got merged

Late-summer commentary merged two papers into one story and put Meta’s name on an 87% score that Meta never reported. There are two papers, and they measure different things.

Meta’s RADAR (Risk Aware Diff Auto Review) is described in “Automating Low-Risk Code Review at Meta: RADAR, Risk Calibration, and Review Efficiency”, submitted to arXiv on May 28, 2026 and revised Jun 12 (v2, still the latest on Oct 8). Thirty-one authors, most at Meta, plus Audris Mockus, Peter Rigby and Nachiappan Nagappan. RADAR is a funnel: it classifies each diff by authorship and source (bot codemod, AI codemod, runbook, human), applies eligibility gates and static heuristics, scores it with a machine-learned Diff Risk Score, runs one LLM reviewer, then deterministic validation before anything lands. The abstract says RADAR “has reviewed 535K+ diffs and landed 331K+” and cut “median diff review wall time by 35%.”

arXiv abstract page for 2605.30208, Automating Low-Risk Code Review at Meta: RADAR, Risk Calibration, and Review Efficiency, listing 31 authors and the abstract with the 535K+ diffs reviewed and 35% review wall time figures Screenshot: arXiv, “[2605.30208] Automating Low-Risk Code Review at Meta: RADAR, Risk Calibration, and Review Efficiency” (submitted May 28, 2026; revised Jun 12, 2026), captured Oct 7, 2026.

What the RADAR paper does not say matters as much. It gives no measurement window (“we cannot provide exact dates”). Its revert rate at about one third and production-incident rate at about one fiftieth of non-RADAR diffs are observational comparisons, not causal ones, as the full paper says.

The abstract also claims a percentage cut in time to close that is larger than 100% and comes with no baseline, so it cannot be a real reduction and is not used here. And RADAR has no second reviewer, no adversary and no critic.

The 87% belongs to “Adversarial Review: Structured Disagreement for Grounded Agentic Code Review” by Eric S. Qiu (Cornell) and Joyce Gill (Stanford), submitted Aug 16, 2026 and listed as accepted to an ICML 2026 workshop (DL4C), where it also appeared as a poster. Three agents: a coder, a reviewer, and a critic that audits the review before the coder edits. On 105 LiveCodeBench tasks, with every call on Claude Sonnet 4.5 and single runs, it scored 87%, against 82% for a five-agent ensemble (MARS, per the paper’s Table 1), 77% zero-shot, 77% for a single reviewer, and 75% for two independent reviewers whose comments the coder applied together.

arXiv abstract page for 2608.18167, Adversarial Review: Structured Disagreement for Grounded Agentic Code Review by Eric S. Qiu and Joyce Gill, showing the abstract and the ICML 2026 workshop acceptance note Screenshot: arXiv, “[2608.18167] Adversarial Review: Structured Disagreement for Grounded Agentic Code Review” (submitted Aug 16, 2026), captured Oct 7, 2026.

Two caveats travel with that number. It is a two-author preprint on a benchmark, not production telemetry. And the critic setup costs more, not less: the full paper reports roughly 4.5 times zero-shot token use on SWE-bench Verified, and claims a place “on the cost-quality Pareto frontier” (three agents beating five), not a cheaper review. Its most useful finding for operators is a failure: without a structured verdict, reviewer and critic drifted into “false consensus,” agreeing without evidence, until the critic was forced to type its disagreement.

Why agent pull requests outgrew one reviewer per PR

RADAR’s abstract states the pressure plainly: at Meta, lines of code per human-landed diff grew 105.9% year over year, diffs per developer rose 51%, and agentic AI accounts for over 80% of that growth, while the share of diffs getting timely review fell. Your repo is smaller, and the curve has the same shape. Agents open PRs at the speed of a loop; reviewers read at the speed of a person.

Adding reviewers scales the wrong side of that equation. Routing scales the right one: most agent PRs are docs, tests and one-file fixes that do not need a model reading them twice, and a few touch auth or money and need a human no matter how good the bots are. Who is allowed to merge is a separate decision, already settled in the PR review agent policy: agents propose, humans merge. This protocol decides who reads.

The risk-routed review protocol, step by step

Step 1: Score every PR before any reviewer sees it

A risk score is a function of facts you already have: which paths the diff touches, how many files, who or what authored it, and whether CI passed. RADAR uses a learned percentile score; you can start with rules and a percentile later. Write the rules down where the routing job can read them. An illustrative starting point, saved as review-routing.yaml:

source_lanes:
  deterministic_codemod: blanket_path
  ai_codemod: score_and_route
  runbook: score_and_route
  human_plus_agent: score_and_route
tiers:
  T0_docs_tests:
    match_all_paths: ["docs/**", "**/*.md", "tests/**", "**/*_test.go"]
  T3_sensitive:
    match_any_path: ["services/auth/**", "services/billing/**", "db/migrations/**", "infra/iam/**"]
  T2_multi_file:
    min_files_changed: 2
  T1_single_file:
    max_files_changed: 1
order: [T3_sensitive, T0_docs_tests, T2_multi_file, T1_single_file]
on_ci_failure: route_to_human
on_extra_reviewers_requested: route_to_human

RADAR sorts diffs by source before it scores them, so the file does too: a deterministic codemod takes one blanket path, and everything else is scored. Order matters: the sensitive tier wins over every other match, so a one-line change in a migration never rides the single-file lane. Copy RADAR’s hard exclusions too. The paper never auto-lands compliance-scoped code, open-source repositories, or diffs that request extra reviewers; neither should you.

Then decide where the cut sits, and move it on evidence. RADAR’s abstract reports that relaxing the Diff Risk Score threshold from the 25th to the 50th percentile raised the approve rate to 60.31%. That is the lever: a looser cut sends more PRs down the cheap lanes, and only your revert and incident rates can tell you whether it was safe. Start strict, log every tier decision, and loosen one percentile band at a time, never during a week when the agents are touching new parts of the repo.

Step 2: Fill in the protocol table for your repo

This is the artifact. The rows are risk tiers; the columns say who reviews, what the critic may do, when a person takes over, what the PR may spend and what it leaves behind. The example below is filled for an illustrative repo: a SaaS monorepo with a web app, a billing service, an auth service and a docs site, receiving about 400 agent PRs a month. Token budgets are illustrative.

Risk tier Example diff (illustrative repo) Who reviews What the critic may do Escalates to a human when Per-PR token budget (illustrative) Evidence left
T0 · docs or tests only README fix; new unit test in web/ None; deterministic checks (lint, tests, link check) Not invoked Any path outside docs/tests appears on re-push; re-score 0 Risk score, CI result
T1 · single-file logic One-function fix in web/src/cart.ts One AI reviewer Not invoked Reviewer confidence below 8/10, any blocking comment, CI failure 40K Risk score, review verdict
T2 · multi-file Refactor across web/ and a shared lib, 6 files AI reviewer + critic Score the review, flag missed risks, hold auto-land; never edits or approves code Disagreement unresolved after 5 rounds, or any DISAGREE_CONCERN the reviewer cannot answer with code 40K reviewer + 45K critic Risk score, review verdict, critic verdict
T3 · auth, payments, migrations New column migration in db/migrations/ plus billing read path Human required; AI reviewer + critic prepare the brief Flag missed risks to the human reviewer; no hold or release rights Always; the human is the gate 60K reviewer + 60K critic Risk score, both verdicts, human approval

The critic grades the review, never the code. It does not leave comments on the diff, does not approve and cannot merge. Its one power, in T2, is to keep a PR from auto-landing until the disagreement is resolved or a person looks.

Step 3: Give the critic a verdict enum, not a voice

A critic that writes free-form prose becomes a second reviewer with a different name. The Adversarial Review paper’s fix was structure: the critic must classify its stance, and every disagreement must point at something. Use three values:

  • AGREE: the review’s claims hold against the code.
  • DISAGREE_EVIDENCE: a review claim is contradicted by specific code, cited by file and line.
  • DISAGREE_CONCERN: an objection without a code citation, such as a risk the review did not mention.

The reviewer must answer each disagreement with code, not confidence. A verdict record the routing job can parse might look like this (illustrative):

{
  "pr": 4182,
  "tier": "T2_multi_file",
  "risk_score": 0.71,
  "review_verdict": "approve_with_comments",
  "critic_verdict": "DISAGREE_EVIDENCE",
  "critic_cites": ["web/src/session.ts:88"],
  "critic_note": "Review says token refresh is unchanged; line 88 now skips refresh on 401.",
  "round": 2,
  "max_rounds": 5
}

Step 4: Freeze the diff while reviewer and critic argue

The paper’s protocol keeps the exchange text-only and edits between rounds, never during one. Copy both rules. If the coding agent pushes while the critic is mid-verdict, the critic is grading a review of code that no longer exists. Pin the commit SHA in the verdict record, cap the exchange at five rounds, and let the coder edit only after a round closes.

Step 5: Spend tokens by tier, and put the critic on the cheaper model

Here is the budget rule: spend the second slot on a critic, and try the critic on the cheaper model, before you buy a second reviewer on the expensive one. Be precise about what supports it. The paper ran every role on the same model and found a reviewer plus critic (87%) beat two independent reviewers (75%). It did not test a cheaper critic.

The cheap-critic half is an inference: grading a review against cited lines is a narrower job than reading a whole diff, so it is the first role to try on a smaller model, and the verdict enum makes a weak critic visible quickly.

Test it on your own PRs before you trust it: for two weeks, run the cheap critic and an expensive critic side by side on the same T2 PRs, and keep the cheap one only if the two disagree on few verdicts and the cheap one’s misses are not concentrated in a single directory.

The worked numbers, all illustrative: 400 agent PRs a month split 120 T0, 160 T1, 90 T2 and 30 T3. Price the expensive reviewer model at $10 per million tokens and the cheap critic model at $2 per million, blended. Then:

  • T1: 160 PRs × 40K tokens on the expensive model = 6.4M tokens, $64.
  • T2: reviewer 40K expensive ($0.40) plus critic 45K cheap ($0.09) = $0.49 a PR; 90 PRs, $44.10. Two expensive reviewers would cost $0.80 a PR, $72.
  • T3: reviewer 60K ($0.60) plus critic 60K ($0.12) = $0.72 a PR; 30 PRs, $21.60, plus the human.

The routed month comes to $129.70 in illustrative model spend. Running two expensive reviewers on every non-docs PR would cost $236 and, if the paper’s result carries over, review worse. The point is less the dollar gap than where the money goes: the critic budget exists only on the 120 PRs that can do real damage.

Horizontal bar chart of the Adversarial Review paper’s LiveCodeBench pass rates on 105 tasks: reviewer plus critic 87%, five-agent ensemble 82%, single reviewer 77%, two independent reviewers 75%, with a stat tile showing Meta RADAR’s 35% lower median diff review wall timeHorizontal bar chart of the Adversarial Review paper’s LiveCodeBench pass rates on 105 tasks: reviewer plus critic 87%, five-agent ensemble 82%, single reviewer 77%, two independent reviewers 75%, with a stat tile showing Meta RADAR’s 35% lower median diff review wall time Qiu and Gill’s Table 1 (105 LiveCodeBench tasks, Claude Sonnet 4.5, single runs) beside RADAR’s 35% review-time figure. Different papers, different measurements; neither is your repo.

Step 6: Write the escalation path before the first disagreement

Disagreements are the protocol working, so decide in advance where each one goes. An illustrative path:

T2 PR, critic returns DISAGREE_EVIDENCE
  -> reviewer answers with code citation
  -> critic re-grades (round n+1, same SHA)
  -> AGREE                          : auto-land eligible after veto window
  -> still DISAGREE after round 5   : assign human, attach both verdicts
T2 PR, critic returns DISAGREE_CONCERN, reviewer cannot cite code
  -> assign human immediately
Any tier, CI fails or path re-scores to T3
  -> assign human, drop pending verdicts

RADAR auto-lands only after a human-veto delay, a window in which any engineer can stop the land. Keep one. A PR that clears both agents still waits long enough for a person to say no.

Step 7: Sample the PRs where both agents said “fine”

False consensus does not announce itself; it looks like a quiet approval. Each week, pull a sample of T2 PRs where the reviewer approved and the critic returned AGREE in round one, and have a person read them cold. In the paper, the failures were a critic that gave way to a confident rebuttal with no evidence behind it, and agreement on hedged comments that turned out to be fabricated. If your sample finds either, tighten the enum’s citation rule before you widen any tier.

Step 8: Keep the evidence where the merge gate can read it

Every PR leaves its risk score, and every reviewed PR leaves its verdicts, keyed by commit SHA. That record is what an overnight merge gate checks before it lets agent work through, and the gate, not this protocol, defines done. Where the review bot runs, which secrets it holds and what its verdict can and cannot prove are covered in the runner and evidence guide; this protocol only adds a second verdict field and a tier label to the record it already keeps.

Flow diagram of risk-routed AI code review: a pull request gets a risk score, is routed to a tier, read by one AI reviewer, audited by a critic on higher tiers, produces a verdict record, and reaches a human gateFlow diagram of risk-routed AI code review: a pull request gets a risk score, is routed to a tier, read by one AI reviewer, audited by a critic on higher tiers, produces a verdict record, and reaches a human gate Seven stops. The critic only appears on the tiers that can afford it, and the human gate is always last.

When the routing table lies: review failure signals

What breaks Signal you would see First action
Tier rules drift behind the repo A new sensitive directory (say services/payouts/) shows T1 PRs in the verdict log Add the path to T3; re-score open PRs touching it
The critic turns into a second reviewer Critic notes comment on style or naming, with no AGREE/DISAGREE value or no citation Reject verdicts without an enum value; re-prompt with the review as the only object
False consensus at T2 Round-one AGREE rate climbs while the weekly cold-read sample finds missed defects Require a cited line for every AGREE on T2 for two weeks; compare sample results
Critic budget leaks into low tiers Critic tokens appear on T0 or T1 PRs in the spend report Fix the router’s tier check; cap critic calls per tier
Evidence keyed to the wrong commit Verdict SHA differs from the merged SHA Block auto-land on SHA mismatch; re-run review on the final commit
Humans rubber-stamp T3 T3 approval minutes fall toward zero while T3 volume rises Put both verdicts at the top of the T3 brief; split large T3 PRs

Review topology across a fleet of coding agents

Once several coding agents share a repo, review stops being a per-PR choice and becomes fleet configuration: one routing file, one verdict schema, one place where tiers and spend are visible. That is the argument for running agents from a command center instead of one terminal each. A critic lane is also one more consumer of provider keys, so it belongs in the inventory the next time you run a fleet key-migration drill.

Keep the lane boundaries clean. This protocol decides who reads a PR and what they may say. Testing whether an agent integration holds up under deliberate attack is a different job with its own plan, set out in the MCP integration attack checklist. Routing review well does not make that testing optional.

FAQ

Are two AI code reviewers better than one?

Not in the one controlled test available. In Qiu and Gill’s August 2026 preprint, two independent reviewers scored 75% on 105 LiveCodeBench tasks, below a single reviewer at 77%. A reviewer plus a critic that audited the review scored 87%. Single runs, one model, one benchmark: test the pattern on your own pull requests.

Does Meta’s RADAR use adversarial review agents?

No. RADAR, described in arXiv paper 2605.30208, runs one LLM reviewer inside a risk-scored funnel with deterministic checks; low-risk diffs can auto-land, everything else goes to a human. The 87% adversarial-review result comes from a separate Cornell and Stanford paper, Adversarial Review, published in August 2026.

What should an AI reviewer critic be allowed to do?

Grade the review, not the code. Give it a fixed verdict (agree, disagree with cited code, or disagree with an uncited concern), let it flag risks the review missed, and let it hold a pull request from auto-landing. It should never comment on the diff, approve or merge anything.

Sources

YOU'RE THROUGH THIS ONE.

Keep connecting the dots.

Back to the library