Shared benchmark scope · OSWorld 2.0

Same-benchmark agent comparison

The entries below share the same benchmark scope. Their exact scaffolds, model versions, dates, and original source values remain visible.

A higher source value only means better performance on this benchmark and configuration. It is not a cross-benchmark or production-readiness score.

Compare like with like

Decision overview

Only configurations from the same benchmark version and task scope can be placed in one comparison.

Source-reported data; missing values remain unknown.
AnthropicOSWorld Agent · Claude Sonnet 4.6 · Max · Standard
Binary completion8.3%Cost per successful taskNot publishedPublished result rangeNot published
Agent details
AnthropicOSWorld Agent · Claude Opus 4.8 · Max · Batched tool
Binary completion20.6%Cost per successful taskNot publishedPublished result rangeNot published
Agent details
15 identical or entirely unreported fields hidden
Agent systemOSWorld Agent · Claude Sonnet 4.6 · Max · StandardOSWorld Agent · Claude Opus 4.8 · Max · Batched tool
Configuration
Base modelClaude Sonnet 4.6Claude Opus 4.8
Source model ID or labelClaude Sonnet 4.6Claude Opus 4.8
ToolsStandard computer-use toolBatched computer-use tool
Outcome
Binary completion8.3%20.6%
Partial score41.5%54.8%
Reported benchmark cost$2,410Not published
Source
Open sourceOpen sourceOpen source

Configuration

OSWorld Agent · Claude Sonnet 4.6 · Max · Standard

Base model
Claude Sonnet 4.6
Source model ID or label
Claude Sonnet 4.6
Tools
Standard computer-use tool

OSWorld Agent · Claude Opus 4.8 · Max · Batched tool

Base model
Claude Opus 4.8
Source model ID or label
Claude Opus 4.8
Tools
Batched computer-use tool

Outcome

OSWorld Agent · Claude Sonnet 4.6 · Max · Standard

Binary completion
8.3%
Partial score
41.5%
Reported benchmark cost
$2,410

OSWorld Agent · Claude Opus 4.8 · Max · Batched tool

Binary completion
20.6%
Partial score
54.8%
Reported benchmark cost
Not published

Source

Source-reported data; missing values remain unknown.

Validate before adoption

Public rankings create a shortlist. A small private trial decides whether the configuration fits your work.

  1. 1

    Choose 10–20 representative tasks and define a clear pass condition.

  2. 2

    Lock the agent, model, tools, permissions, and budget before testing.

  3. 3

    Run each important task more than once; record success, cost, time, and takeovers.

  4. 4

    Review failures and data-handling risks before a production decision.

Agent directory and evaluations