Shared benchmark scope · OSWorld 2.0

Same-benchmark agent comparison

The entries below share the same benchmark scope. Their exact scaffolds, model versions, dates, and original source values remain visible.

A higher source value only means better performance on this benchmark and configuration. It is not a cross-benchmark or production-readiness score.

Compare like with like

Decision overview

Only configurations from the same benchmark version and task scope can be placed in one comparison.

Source-reported data; missing values remain unknown.
AnthropicOSWorld Agent · Claude Opus 4.7 · Max · Standard
Binary completion13.9%Cost per successful taskNot publishedPublished result rangeNot published
Agent details
AnthropicOSWorld Agent · Claude Opus 4.8 · Max · Batched tool
Binary completion20.6%Cost per successful taskNot publishedPublished result rangeNot published
Agent details
15 identical or entirely unreported fields hidden
Agent systemOSWorld Agent · Claude Opus 4.7 · Max · StandardOSWorld Agent · Claude Opus 4.8 · Max · Batched tool
Configuration
Base modelClaude Opus 4.7Claude Opus 4.8
Source model ID or labelClaude Opus 4.7Claude Opus 4.8
ToolsStandard computer-use toolBatched computer-use tool
Outcome
Binary completion13.9%20.6%
Partial score49.1%54.8%
Reported benchmark cost$3,870Not published
Source
Open sourceOpen sourceOpen source

Configuration

OSWorld Agent · Claude Opus 4.7 · Max · Standard

Base model
Claude Opus 4.7
Source model ID or label
Claude Opus 4.7
Tools
Standard computer-use tool

OSWorld Agent · Claude Opus 4.8 · Max · Batched tool

Base model
Claude Opus 4.8
Source model ID or label
Claude Opus 4.8
Tools
Batched computer-use tool

Outcome

OSWorld Agent · Claude Opus 4.7 · Max · Standard

Binary completion
13.9%
Partial score
49.1%
Reported benchmark cost
$3,870

OSWorld Agent · Claude Opus 4.8 · Max · Batched tool

Binary completion
20.6%
Partial score
54.8%
Reported benchmark cost
Not published

Source

Source-reported data; missing values remain unknown.

Validate before adoption

Public rankings create a shortlist. A small private trial decides whether the configuration fits your work.

  1. 1

    Choose 10–20 representative tasks and define a clear pass condition.

  2. 2

    Lock the agent, model, tools, permissions, and budget before testing.

  3. 3

    Run each important task more than once; record success, cost, time, and takeovers.

  4. 4

    Review failures and data-handling risks before a production decision.

Agent directory and evaluations