Shared benchmark scope · OSWorld 2.0
Same-benchmark agent comparison
The entries below share the same benchmark scope. Their exact scaffolds, model versions, dates, and original source values remain visible.
A higher source value only means better performance on this benchmark and configuration. It is not a cross-benchmark or production-readiness score.
Compare like with like
Decision overview
Only configurations from the same benchmark version and task scope can be placed in one comparison.
| Agent system | OSWorld Agent · MiniMax M3 · Enabled · Standard | OSWorld Agent · Claude Opus 4.8 · Max · Batched tool |
|---|---|---|
| Configuration | ||
| Base model | MiniMax M3 | Claude Opus 4.8 |
| Source model ID or label | MiniMax M3 | Claude Opus 4.8 |
| Tools | Standard computer-use tool | Batched computer-use tool |
| Outcome | ||
| Binary completion | 4.6% | 20.6% |
| Partial score | 22.3% | 54.8% |
| Reported benchmark cost | $258.78 | Not published |
| Source | ||
| Open source | Open source ↗ | Open source ↗ |
Configuration
OSWorld Agent · MiniMax M3 · Enabled · Standard
- Base model
- MiniMax M3
- Source model ID or label
- MiniMax M3
- Tools
- Standard computer-use tool
OSWorld Agent · Claude Opus 4.8 · Max · Batched tool
- Base model
- Claude Opus 4.8
- Source model ID or label
- Claude Opus 4.8
- Tools
- Batched computer-use tool
Outcome
OSWorld Agent · MiniMax M3 · Enabled · Standard
- Binary completion
- 4.6%
- Partial score
- 22.3%
- Reported benchmark cost
- $258.78
OSWorld Agent · Claude Opus 4.8 · Max · Batched tool
- Binary completion
- 20.6%
- Partial score
- 54.8%
- Reported benchmark cost
- Not published
Source
OSWorld Agent · MiniMax M3 · Enabled · Standard
OSWorld Agent · Claude Opus 4.8 · Max · Batched tool
Source-reported data; missing values remain unknown.
Validate before adoption
Public rankings create a shortlist. A small private trial decides whether the configuration fits your work.
- 1
Choose 10–20 representative tasks and define a clear pass condition.
- 2
Lock the agent, model, tools, permissions, and budget before testing.
- 3
Run each important task more than once; record success, cost, time, and takeovers.
- 4
Review failures and data-handling risks before a production decision.