Agent benchmarks

OSWorld 2.0 agent benchmark

OpenGPT keeps 10 published rows attached to the exact OSWorld 2.0 benchmark version and environment shown below.

Tasks
108
Published rows
10
Source data date
Identity reviewed
Current

Evidence scope

task version v2026.06.24 · 500-step budget. long-horizon computer workflows. Desktop applications and operating-system workflows in reproducible VMs. 500 steps.

How OpenGPT handles this evidence

OpenGPT preserves the publisher's metric, task scope, environment, version, configuration fields, and source date. It does not create a cross-benchmark overall Agent score.

Comparison boundaries

  • A higher result applies only to the displayed benchmark version and conditions.
  • Agent results reflect the full system, not the base model alone.
  • Missing cost, reliability, safety, or latency data remains unreported rather than being inferred.
  • Rows preserve the complete agent, model, tool mode, task version, step budget, and source-reported metrics.

task version v2026.06.24 · 500-step budget

Showing 10 of 10 published rows. Every result remains attached to the source benchmark and configuration.

  1. OSWorld Agent · Claude Opus 4.8 · Max · Batched toolAnthropic · Claude Opus 4.8#1Binary completion: 20.6%OSWorld 2.0 computer-use agent
  2. OSWorld Agent · Claude Opus 4.8 · Max · StandardAnthropic · Claude Opus 4.8#2Binary completion: 18.52%OSWorld 2.0 computer-use agent
  3. OSWorld Agent · Claude Opus 4.7 · Max · Batched toolAnthropic · Claude Opus 4.7#3Binary completion: 18.2%OSWorld 2.0 computer-use agent
  4. OSWorld Agent · Claude Opus 4.7 · Max · StandardAnthropic · Claude Opus 4.7#4Binary completion: 13.9%OSWorld 2.0 computer-use agent
  5. OSWorld Agent · GPT-5.5 · Xhigh · Batch toolOpenAI · GPT-5.5#5Binary completion: 13.0%OSWorld 2.0 computer-use agent
  6. OSWorld Agent · Claude Sonnet 4.6 · Medium · StandardAnthropic · Claude Sonnet 4.6#6Binary completion: 9.3%OSWorld 2.0 computer-use agent
  7. OSWorld Agent · Claude Sonnet 4.6 · Max · StandardAnthropic · Claude Sonnet 4.6#7Binary completion: 8.3%OSWorld 2.0 computer-use agent
  8. OSWorld Agent · MiniMax M3 · Enabled · StandardMiniMax · MiniMax M3#8Binary completion: 4.6%OSWorld 2.0 computer-use agent
  9. OSWorld Agent · Kimi 2.6 · Enabled · StandardMoonshot AI · Kimi 2.6#8Binary completion: 4.6%OSWorld 2.0 computer-use agent
  10. OSWorld Agent · Qwen 3.7-Plus · Thinking · StandardAlibaba Qwen · Qwen 3.7-Plus#10Binary completion: 2.8%OSWorld 2.0 computer-use agent