Agent configuration · OSWorld 2.0

OSWorld Agent · MiniMax M3 · Enabled · Standard

The ranked object: model, scaffold, tools, permissions, budget, and environment together.

Binary completion4.6%
Cost per successful taskNot published
Published result rangeNot published

01

Configuration

Agent systemOSWorld Agent · MiniMax M3 · Enabled · Standard
Base modelMiniMax M3
Source model ID or labelMiniMax M3
ScaffoldOSWorld 2.0 computer-use agent
Scaffold versionOSWorld 2.0
ToolsStandard computer-use tool
EnvironmentDesktop and operating-system workflows in reproducible virtual machines
Budget500 steps

02

Outcome

Binary completion4.6%
Partial score22.3%
Cost per successful taskNot published
Reported benchmark cost$258.78
Median completion timeNot published
Human interventionNot published

03

Operation

Repeat-run evidenceNot published
Published result rangeNot published
Safety noteNot published
Comparability grouposworld-2-v2026-06-24-500
Evaluation date
Data captured

04

Trust

EvidenceListed by the benchmark publisher
SourceOSWorld 2.0
Source snapshotTask version v2026.06.24 · 500-step budget
Source data dateNot published by source
Retrieved by OpenGPT
Ranking last changedNo comparable history yet
Last checked
Tasks108 long-horizon computer workflows
Evidence levelPublisher maintained
Evidence note

Rows preserve the reported agent, model, tool mode, task version, step limit, and source metrics.

Open audit

What this result does not prove

This result applies to the exact system and source conditions shown here. It does not establish a universal best agent, production reliability, data governance, or performance on a different benchmark version.

Compare like with like

Only configurations from the same benchmark version and task scope can be placed in one comparison.

Source-reported data; missing values remain unknown.

Validate before adoption

Public rankings create a shortlist. A small private trial decides whether the configuration fits your work.

  1. 1

    Choose 10–20 representative tasks and define a clear pass condition.

  2. 2

    Lock the agent, model, tools, permissions, and budget before testing.

  3. 3

    Run each important task more than once; record success, cost, time, and takeovers.

  4. 4

    Review failures and data-handling risks before a production decision.

All AI Agent rankings