Review the Trajectory, Not the Reel: OSWorld 2.0's Checkpoints for Your Own Computer-Use Runs
Three OSWorld 2.0 numbers went around in September and no two measure the same thing. Borrow its checkpoint grader to review your own computer-use runs instead.
Go deeper. Build your own.
81.8%, 72.6% and 44.33%. Those are three OSWorld 2.0 figures from the past five weeks, and no two of them measure the same thing. Anthropic’s is a partial score on a task release vendors now call “OSWorld 2.1.” OpenAI’s is a partial score on an offline subset. The 44.33% is binary completion at 500 steps from the top row of Snorkel AI’s leaderboard, and that same row reads 77.67% partial.
Put them in one ranked table and you have a reel: footage that looks like evidence and isn’t.
The part of OSWorld 2.0 worth taking home is the grader. Each task carries about 27 weighted checkpoints, almost all of them functional checks on the end state of the machine, plus audits for reward hacking and false negatives, separate safety reports, and a simulated user the agent can ask when something is missing. That design ports straight onto the computer-use jobs you already run. Grade your own runs on end state and process, and let vendor numbers into your notes only with their labels attached.
OSWorld 2.0, its three task releases, and the names vendors use
OSWorld 2.0 comes from the XLANG Lab at the University of Hong Kong, the group behind the original OSWorld. The repo, xlang-ai/OSWorld-V2, dates the release Jun 26, 2026, and the paper, “OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks”, went up on arXiv two days later. The numbers explain why a second version was needed: 108 tasks across 31 self-hosted web environments and desktop apps, a median of about 1.6 hours for a human, and an average of 318 tool calls per task with Claude Opus 4.7 at maximum thinking, “compared with about 30 in OSWorld 1.0.” The paper puts frontier agents at 79–83% binary on OSWorld-Verified, the cleaned-up first version. That board is close to saturated.
Screenshot: arXiv, “[2606.29537] OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks” (submitted Jun 28, 2026; revised Jul 13, 2026), captured Oct 5, 2026.
At publication the best result was Claude Opus 4.8 with maximum thinking and batched tool calls: 20.6% of tasks completed, at a 54.8% partial score, under the paper’s “primary binary-completion metric at 500 steps.” Only 11.53% of the total score comes from model-based judging; the rest is functional checks on environment state.
Snorkel AI did not build it. Snorkel funds it through its Open Benchmarks Grants, whose first wave (announced Jul 24) included OSWorld 2.0, and it hosts a leaderboard for it. If you have seen the benchmark called “OSWorld-Pro” or “Snorkel 2.0,” neither name exists.
Screenshot: Snorkel AI, “OSWorld 2.0: Long-Horizon Computer-Use Benchmark” (leaderboard last updated Oct 2, 2026), captured Oct 5, 2026.
The release names matter more than the version number. The original task set is osworld-v2-2026.06.24 (its hosted sites were retired Sep 8), an update shipped as osworld-v2-2026.08.08, and on Sep 16 the lab released osworld-v2.1, a bug-fix release with updated agents that is now the recommended set. So when Anthropic’s Claude Opus 5.5 launch on Sep 22 reported “OSWorld 2.1: 81.8% partial,” that is the v2.1 task release of OSWorld 2.0, not a new benchmark. The four figures in circulation, with what each source states:
- Anthropic, Claude Opus 5.5 (Sep 22): 81.8% partial on v2.1, with “adaptive thinking at max effort” per the page’s footnote, evaluated with production safeguards enabled. No step budget is given.
- OpenAI, GPT-6 Astra (launched early September): 72.6%, labeled “OSWorld 2.0 (v2026.08.08, offline set, partial score),” at roughly 40 minutes per task.
- Anthropic, Claude Fable 5.1 (early September): 41.7% strict and 77.9% partial “on the benchmark authors’ August 2026 task release; Fable 5 and Opus 5 were re-run under the same conditions.”
- Snorkel’s board (updated Oct 2): the top binary row at 500 steps is Claude Opus 5 at max effort, 44.33% binary and 77.67% partial. The same model appears seven times at different effort levels, with binary scores from 28.98% to 44.33%. The page does not say which task release those runs used or who ran them.
Why a leaderboard number can’t sign off a computer-use job
A leaderboard figure describes a model on someone else’s 108 tasks, under someone else’s step budget and effort setting, graded by someone else’s checkers. Your question is narrower and harder: did this agent do this job correctly this time, and would you know if it hadn’t? Reading benchmarks in general is its own skill; this piece is about the review you run after your own jobs.
OSWorld 2.0’s maintainers already model the habit. The repo says verified-board results require running the agent code with the maintainers or sharing “monitoring data and trajectories,” and it is explicit: “If you want your results to be verified and displayed on the verified leaderboard, you need to schedule a meeting with us…” Even the benchmark’s authors won’t take a number without the trajectory. Neither should you.
Build a trajectory-review sheet for your computer-use runs
The sheet has two halves. The top half grades state: weighted checkpoints, logged as binary and partial. The bottom half grades process: four questions about how the agent got there. Fill one per run, keep it next to the run’s transcript, and review the week’s sheets together.
The trajectory-review sheet: state first (checkpoints, binary, partial), process second (four questions), one ledger row per run.
Step 1: Pick one job and freeze its end state before the run
Start with a job that already reaches the screen in your desk-job inventory. Write its end state in terms a script or a second person can check, and write it before the run. A finish line written after reading the transcript bends toward whatever the agent did.
This is not the same as a finish line the transcript can prove, which works when the transcript is the only evidence. OSWorld 2.0 grades the environment, not the narration, and computer-use jobs need that: the agent’s account of what it clicked is the least reliable artifact the run produces.
Step 2: Write 5–10 weighted checkpoints on end state
OSWorld 2.0 averages 27.25 checkpoints per task because its tasks are hours long. Your jobs are shorter, so 5–10 is enough. Each checkpoint is a functional check on state: a file exists with a hash, a record ID changed, a form was submitted with N fields, a setting reads back a value. Weight by consequence, and keep anything a model has to judge to a small share of the total, the way the paper holds model judging to 11.53% of its score.
Here is an illustrative sheet for one job, re-keying 30 line items from scanned delivery notes into a legacy desktop ERP:
| # | Checkpoint (end state) | How you check it | Weight |
|---|---|---|---|
| 1 | All 30 source line items exist as ERP records in the batch | Count query by batch reference | 3 |
| 2 | Quantities match the source notes | Diff 10 sampled records against the scans | 3 |
| 3 | No records created outside the batch | Record count delta for the day equals 30 | 2 |
| 4 | Batch status is “pending review”, not “posted” | Read the status field back | 2 |
| 5 | Supplier ID matches the delivery note header | Field read-back | 1 |
| 6 | Source scans moved to the processed folder, hashes unchanged | Before/after hash list | 1 |
| 7 | Run log lists the created record IDs | Compare log to the count query | 1 |
Checkpoint 4 carries a constraint from the opening instruction. You will need it again in Step 4.
Step 3: Log binary and partial on every run
Binary asks whether every checkpoint passed. Partial is the passed weight divided by total weight. Log both, because they answer different questions: partial tells you how far the agent got, binary tells you whether you can stop checking. Snorkel’s top row is the clearest illustration you’ll get, one run that reads 44.33% binary and 77.67% partial.
The arithmetic is small enough to keep beside the sheet:
checks = [("records exist", 3, True), ("quantities match", 3, True),
("no stray records", 2, True), ("status pending", 2, False),
("supplier id", 1, True), ("scans unchanged", 1, True),
("log lists ids", 1, True)]
total = sum(w for _, w, _ in checks)
partial = sum(w for _, w, ok in checks if ok) / total # 11 / 13 = 0.85
binary = all(ok for _, _, ok in checks) # False
That run scores 85% partial and fails. It also posted records you told it to leave pending, which is exactly the kind of failure a partial score flatters. Use binary for any decision that removes a human from the loop; use partial to see where runs stall.
Step 4: Ask four process questions of every trajectory
State checks tell you what happened. Process questions tell you whether it will happen again. The OSWorld 2.0 abstract names the pattern in one line: current agents “lose track of constraints, miss information that arrives mid-task, guess rather than ask the user, and skip verification.” Snorkel’s Sep 3 write-up by Mengqi Yuan sorts failures into four modes (information missing, perception, verification, memory), and the paper’s safety reports cover credential leaks, UI bypassing and resource exhaustion. Turn those into four questions:
| Question | Where it comes from | What to look for in the trajectory | A “no” means |
|---|---|---|---|
| Did it ask when information was missing? | Information-missing mode; simulated user | A field filled with no source, an assumption stated as fact | Tag “information”; add the missing input to the job spec |
| Did it verify before claiming done? | Verification mode | “Done” with no read-back, no re-open of the saved file | Tag “verification”; add a read-back step to the prompt |
| Did it stay in the UI and away from credentials? | Safety reports; reward-hacking audits | Calls to APIs or files outside the approved surface, password stores opened, a check satisfied by editing its input | Tag “safety”; stop the job until reviewed |
| Did it keep the early constraints? | Memory mode | An instruction from the first turn violated late in the run | Tag “memory”; restate the constraint as a checkpoint |
Perception failures (misread a field, clicked the wrong row) mostly show up as failed checkpoints rather than process answers. Tag them “perception” when the transcript shows the agent acting on something that wasn’t on screen.
Step 5: Give the agent a way to ask, then grade whether it used it
OSWorld 2.0 lets agents query a simulated user, which is what makes “did it ask?” a fair question. Your runs need the same thing: a question channel the agent is told about, whether that is an approval prompt, a chat reply or a ticket comment. Without one, question 1 always reads “not applicable” and the agent’s guesses never show up as a category.
When the agent does ask, record the question and the answer on the sheet. A good question is a signal that the job spec is missing something; three runs asking the same thing means the spec is wrong, not the agent.
Step 6: Run the two audits OSWorld runs on itself
The paper audits its own grader for reward hacking (a checkpoint passed for the wrong reason) and false negatives (the job was done, the checker said no). Copy both, weekly:
- Reward-hacking audit: pick two passed runs and read the trajectory end to end. Look for checks satisfied by editing the thing being checked: a count file rewritten, an export trimmed to match the expected total, a status field set directly through a back door.
- False-negative audit: pick two failed runs and check by hand. If the work was correct and the checker was too strict (a date format, a trailing space in a reference), fix the checker and note the change with a date.
Fix the checker, never the score. A sheet whose checkpoints change every week without dated notes stops meaning anything.
Step 7: Keep a reel normalizer for vendor numbers
Your own sheets cover your jobs. Vendor numbers still arrive, and they will end up in a model decision. Give them an intake rule: a number enters your notes only with metric (binary or partial), task release, step budget, effort and runner (vendor, benchmark authors or independent). Anything missing a column goes in a separate “reel” tab that never feeds a decision. The runner column matters more than it looks: the steel.dev OSWorld 2.0 aggregator warns that its “rows mix benchmark-author runs, vendor self-reports, and independent co-author runs.”
Three OSWorld 2.0 numbers, three different measurements: metric, release, budget, effort and runner as each source states them. They are not a ranking.
Applied to the figures above, as each source stated them when captured on Oct 5, the normalizer looks like this:
| Figure | Metric | Task release | Step budget | Effort | Runner | Usable for a decision? |
|---|---|---|---|---|---|---|
| 81.8%, Claude Opus 5.5 | Partial | osworld-v2.1 | Not stated | Max (adaptive thinking) | Vendor | Reel tab |
| 72.6%, GPT-6 Astra | Partial | 2026.08.08, offline set | Not stated (about 40 min per task) | Not stated | Vendor | Reel tab |
| 41.7% / 77.9%, Claude Fable 5.1 | Binary and partial | August 2026 release | Not stated | Not stated | Vendor, baselines re-run | Closer, still missing two columns |
| 44.33% / 77.67%, Claude Opus 5 | Binary and partial | Not stated | 500 steps | Max | Not stated | Closer, still missing two columns |
Effort alone moves one model’s binary score from 28.98% to 44.33% on Snorkel’s board. A number with no effort label could sit anywhere in that range. When a figure does clear intake, it becomes one input to the evidence that promotes a model, next to your own sheets, never instead of them.
Step 8: Review the week’s sheets on one schedule
Pick a fixed slot, the same way you would for a document-intake bake-off, and run the checklist:
- Every computer-use run from the week has a sheet with binary, partial and four answers.
- Failure tags counted by category: information, perception, verification, memory, safety.
- Two passes audited for reward hacking, two fails audited for false negatives.
- Checkpoint or weight changes carry a dated note and a reason.
- Any job proposed for unattended runs shows consecutive binary passes, not a partial average.
- New vendor figures sorted into the decision tab or the reel tab.
Where trajectory review fails, and the signal for each
Checkpoints that grade the narration. The signal is a checkpoint whose “how you check it” column cites the agent’s own message. Replace it with a read-back of state, or delete it.
Partial-score comfort. The signal is partial rising week over week while binary stays flat. A job at 85% partial for a month is a job that fails every time in a slightly different place.
Weights tuned to a model. The signal is weight changes clustered around a model switch. Freeze weights for the length of any comparison and change them only with a dated note.
“Did it ask?” is always not applicable. The signal is a column of blanks in question 1. The agent has no question channel, so its guesses are invisible. Add the channel before you grade the answer.
The model you graded isn’t the model that ran. Anthropic’s Opus 5.5 footnotes describe evaluation with production safeguards on; where they intervened, cybersecurity tasks were completed by Claude Opus 4.8 and biology and frontier-LLM-development tasks by Claude Opus 5. Your runs can hit the same swap, so log the served model on the sheet and treat a mismatch as a separate failure tag.
Reel creep. The signal is a vendor percentage in a promotion doc without its five labels. Send it back to the reel tab.
Trajectories are fleet evidence, so store them like it
A review sheet is only as good as the trajectory behind it, and trajectories are the first thing to go missing when runs spread across vendors, machines and cloud workers. Keep the transcript, the checkpoint screenshots and the sheet together under one run ID, so the weekly review reads one folder instead of five consoles.
That is the same discipline as replaying any agent incident: you can’t replay what you can’t see. OSWorld 2.0’s maintainers won’t verify a score without monitoring data and trajectories. Hold your own fleet to the same rule.
FAQ
What is OSWorld 2.0?
OSWorld 2.0 is a computer-use benchmark from XLANG Lab at the University of Hong Kong, released June 26, 2026. It has 108 long-horizon tasks across web environments and desktop apps, graded on about 27 weighted end-state checkpoints per task, with binary completion at 500 steps as the headline metric.
Is OSWorld 2.1 a new benchmark?
No. “OSWorld 2.1” in vendor posts refers to osworld-v2.1, a bug-fix task release of OSWorld 2.0 published September 16, 2026, with updated agents. Earlier task releases are dated June 24 and August 8. Always record which release a number came from, since scores across releases are not directly comparable.
What is the difference between binary and partial scores in OSWorld 2.0?
Binary counts a task as done only when every checkpoint passes. Partial is the weighted share of checkpoints passed. The same run can differ sharply: Snorkel’s top row reads 44.33% binary and 77.67% partial. Use binary for decisions that remove a human; use partial to find where runs stall.
Sources
- XLANG Lab, “OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks,” arXiv 2606.29537 (v1 Jun 28, 2026; v2 Jul 13, 2026): https://arxiv.org/abs/2606.29537
- XLANG Lab, OSWorld-V2 repository, releases and verification rules (osworld-v2.1, Sep 16, 2026): https://github.com/xlang-ai/OSWorld-V2
- Snorkel AI, OSWorld 2.0 leaderboard (last updated Oct 2, 2026): https://snorkel.ai/leaderboard/os-world-2-0/
- Snorkel AI blog, Mengqi Yuan, “OSWorld 2.0: Why Long-Horizon Computer-Use Agents Still Fail Four Out of Five Tasks” (Sep 3, 2026): https://snorkel.ai/blog/osworld-2-0-why-computer-use-agents-fail-most-tasks/
- PR Newswire, Snorkel AI first wave of Open Benchmarks Grants projects (Jul 24, 2026): https://www.prnewswire.com/news-releases/snorkel-ai-highlights-first-wave-of-open-benchmarks-grants-projects-302833805.html
- Anthropic, “Introducing Claude Opus 5.5” (Sep 22, 2026): https://www.anthropic.com/claude-opus-5-5
- OpenAI, “GPT-6 Astra” (early Sep 2026; updated Sep 22 and Sep 29): https://openai.com/index/gpt-6-astra/
- Anthropic, “Introducing Claude Fable 5.1 and Claude Mythos 5.1” (early Sep 2026): https://www.anthropic.com/claude-fable-and-mythos-5-1
- Steel.dev, OSWorld 2.0 aggregator leaderboard (secondary; updated Sep 30, 2026): https://leaderboard.steel.dev/leaderboards/osworld-2/
