Shared benchmark scope · OSWorld 2.0

Same-benchmark agent comparison

The entries below share the same benchmark scope. Their exact scaffolds, model versions, dates, and original source values remain visible.

A higher source value only means better performance on this benchmark and configuration. It is not a cross-benchmark or production-readiness score.

Compare like with like

Decision overview

Only configurations from the same benchmark version and task scope can be placed in one comparison.

Source-reported data; missing values remain unknown.
OpenAIOSWorld Agent · GPT-5.5 · Xhigh · Batch tool
Binary completion13.0%Cost per successful taskNot publishedPublished result rangeNot published
Agent details
AnthropicOSWorld Agent · Claude Opus 4.8 · Max · Batched tool
Binary completion20.6%Cost per successful taskNot publishedPublished result rangeNot published
Agent details
15 identical or entirely unreported fields hidden
Agent systemOSWorld Agent · GPT-5.5 · Xhigh · Batch toolOSWorld Agent · Claude Opus 4.8 · Max · Batched tool
Configuration
Base modelGPT-5.5Claude Opus 4.8
Source model ID or labelGPT-5.5Claude Opus 4.8
ToolsBatch computer-use toolBatched computer-use tool
Outcome
Binary completion13.0%20.6%
Partial score49.5%54.8%
Reported benchmark cost$2,750Not published
Source
Open sourceOpen sourceOpen source

Configuration

OSWorld Agent · GPT-5.5 · Xhigh · Batch tool

Base model
GPT-5.5
Source model ID or label
GPT-5.5
Tools
Batch computer-use tool

OSWorld Agent · Claude Opus 4.8 · Max · Batched tool

Base model
Claude Opus 4.8
Source model ID or label
Claude Opus 4.8
Tools
Batched computer-use tool

Outcome

OSWorld Agent · GPT-5.5 · Xhigh · Batch tool

Binary completion
13.0%
Partial score
49.5%
Reported benchmark cost
$2,750

OSWorld Agent · Claude Opus 4.8 · Max · Batched tool

Binary completion
20.6%
Partial score
54.8%
Reported benchmark cost
Not published

Source

Source-reported data; missing values remain unknown.

Validate before adoption

Public rankings create a shortlist. A small private trial decides whether the configuration fits your work.

  1. 1

    Choose 10–20 representative tasks and define a clear pass condition.

  2. 2

    Lock the agent, model, tools, permissions, and budget before testing.

  3. 3

    Run each important task more than once; record success, cost, time, and takeovers.

  4. 4

    Review failures and data-handling risks before a production decision.

Agent directory and evaluations