Shared benchmark scope · OSWorld 2.0

Same-benchmark agent comparison

The entries below share the same benchmark scope. Their exact scaffolds, model versions, dates, and original source values remain visible.

A higher source value only means better performance on this benchmark and configuration. It is not a cross-benchmark or production-readiness score.

Compare like with like

Decision overview

Only configurations from the same benchmark version and task scope can be placed in one comparison.

Source-reported data; missing values remain unknown.
Alibaba QwenOSWorld Agent · Qwen 3.7-Plus · Thinking · Standard
Binary completion2.8%Cost per successful taskNot publishedPublished result rangeNot published
Agent details
AnthropicOSWorld Agent · Claude Opus 4.8 · Max · Batched tool
Binary completion20.6%Cost per successful taskNot publishedPublished result rangeNot published
Agent details
15 identical or entirely unreported fields hidden
Agent systemOSWorld Agent · Qwen 3.7-Plus · Thinking · StandardOSWorld Agent · Claude Opus 4.8 · Max · Batched tool
Configuration
Base modelQwen 3.7-PlusClaude Opus 4.8
Source model ID or labelQwen 3.7-PlusClaude Opus 4.8
ToolsStandard computer-use toolBatched computer-use tool
Outcome
Binary completion2.8%20.6%
Partial score21.5%54.8%
Reported benchmark cost$411.56Not published
Source
Open sourceOpen sourceOpen source

Configuration

OSWorld Agent · Qwen 3.7-Plus · Thinking · Standard

Base model
Qwen 3.7-Plus
Source model ID or label
Qwen 3.7-Plus
Tools
Standard computer-use tool

OSWorld Agent · Claude Opus 4.8 · Max · Batched tool

Base model
Claude Opus 4.8
Source model ID or label
Claude Opus 4.8
Tools
Batched computer-use tool

Outcome

OSWorld Agent · Qwen 3.7-Plus · Thinking · Standard

Binary completion
2.8%
Partial score
21.5%
Reported benchmark cost
$411.56

OSWorld Agent · Claude Opus 4.8 · Max · Batched tool

Binary completion
20.6%
Partial score
54.8%
Reported benchmark cost
Not published

Source

Source-reported data; missing values remain unknown.

Validate before adoption

Public rankings create a shortlist. A small private trial decides whether the configuration fits your work.

  1. 1

    Choose 10–20 representative tasks and define a clear pass condition.

  2. 2

    Lock the agent, model, tools, permissions, and budget before testing.

  3. 3

    Run each important task more than once; record success, cost, time, and takeovers.

  4. 4

    Review failures and data-handling risks before a production decision.

Agent directory and evaluations