Shared benchmark scope · OSWorld 2.0

Same-benchmark agent comparison

The entries below share the same benchmark scope. Their exact scaffolds, model versions, dates, and original source values remain visible.

A higher source value only means better performance on this benchmark and configuration. It is not a cross-benchmark or production-readiness score.

Compare like with like

Decision overview

Only configurations from the same benchmark version and task scope can be placed in one comparison.

Source-reported data; missing values remain unknown.
Moonshot AIOSWorld Agent · Kimi 2.6 · Enabled · Standard
Binary completion4.6%Cost per successful taskNot publishedPublished result rangeNot published
Agent details
AnthropicOSWorld Agent · Claude Opus 4.8 · Max · Batched tool
Binary completion20.6%Cost per successful taskNot publishedPublished result rangeNot published
Agent details
15 identical or entirely unreported fields hidden
Agent systemOSWorld Agent · Kimi 2.6 · Enabled · StandardOSWorld Agent · Claude Opus 4.8 · Max · Batched tool
Configuration
Base modelKimi 2.6Claude Opus 4.8
Source model ID or labelKimi 2.6Claude Opus 4.8
ToolsStandard computer-use toolBatched computer-use tool
Outcome
Binary completion4.6%20.6%
Partial score22.1%54.8%
Reported benchmark cost$708Not published
Source
Open sourceOpen sourceOpen source

Configuration

OSWorld Agent · Kimi 2.6 · Enabled · Standard

Base model
Kimi 2.6
Source model ID or label
Kimi 2.6
Tools
Standard computer-use tool

OSWorld Agent · Claude Opus 4.8 · Max · Batched tool

Base model
Claude Opus 4.8
Source model ID or label
Claude Opus 4.8
Tools
Batched computer-use tool

Outcome

OSWorld Agent · Kimi 2.6 · Enabled · Standard

Binary completion
4.6%
Partial score
22.1%
Reported benchmark cost
$708

OSWorld Agent · Claude Opus 4.8 · Max · Batched tool

Binary completion
20.6%
Partial score
54.8%
Reported benchmark cost
Not published

Source

Source-reported data; missing values remain unknown.

Validate before adoption

Public rankings create a shortlist. A small private trial decides whether the configuration fits your work.

  1. 1

    Choose 10–20 representative tasks and define a clear pass condition.

  2. 2

    Lock the agent, model, tools, permissions, and budget before testing.

  3. 3

    Run each important task more than once; record success, cost, time, and takeovers.

  4. 4

    Review failures and data-handling risks before a production decision.

Agent directory and evaluations