エージェント評価
OSWorld 2.0エージェント評価
OSWorld 2.0の表示版と環境に結び付いた公開行10件を保持します。
- 課題
- 108
- 公開行
- 10
- 情報源データ日
- ID確認日
- 現在
根拠の範囲
task version v2026.06.24 · 500-step budget. long-horizon computer workflows. Desktop applications and operating-system workflows in reproducible VMs. 500 steps.
OpenGPTでの根拠の扱い
公開元の指標、課題範囲、環境、版、構成、日付を保持し、評価を横断した総合点は作りません。
比較できる範囲
- 高い値は表示された評価版と条件にだけ適用されます。
- 結果は基盤モデル単体ではなくエージェント全体を反映します。
- 費用、信頼性、安全性、速度の欠損値は推測しません。
- Rows preserve the complete agent, model, tool mode, task version, step budget, and source-reported metrics.
task version v2026.06.24 · 500-step budget
公開10件のうち10件を表示します。結果は元の評価と構成に結び付いています。
- OSWorld Agent · Claude Opus 4.8 · Max · Batched toolAnthropic · Claude Opus 4.8#1Binary completion: 20.6%OSWorld 2.0 computer-use agent
- OSWorld Agent · Claude Opus 4.8 · Max · StandardAnthropic · Claude Opus 4.8#2Binary completion: 18.52%OSWorld 2.0 computer-use agent
- OSWorld Agent · Claude Opus 4.7 · Max · Batched toolAnthropic · Claude Opus 4.7#3Binary completion: 18.2%OSWorld 2.0 computer-use agent
- OSWorld Agent · Claude Opus 4.7 · Max · StandardAnthropic · Claude Opus 4.7#4Binary completion: 13.9%OSWorld 2.0 computer-use agent
- OSWorld Agent · GPT-5.5 · Xhigh · Batch toolOpenAI · GPT-5.5#5Binary completion: 13.0%OSWorld 2.0 computer-use agent
- OSWorld Agent · Claude Sonnet 4.6 · Medium · StandardAnthropic · Claude Sonnet 4.6#6Binary completion: 9.3%OSWorld 2.0 computer-use agent
- OSWorld Agent · Claude Sonnet 4.6 · Max · StandardAnthropic · Claude Sonnet 4.6#7Binary completion: 8.3%OSWorld 2.0 computer-use agent
- OSWorld Agent · MiniMax M3 · Enabled · StandardMiniMax · MiniMax M3#8Binary completion: 4.6%OSWorld 2.0 computer-use agent
- OSWorld Agent · Kimi 2.6 · Enabled · StandardMoonshot AI · Kimi 2.6#8Binary completion: 4.6%OSWorld 2.0 computer-use agent
- OSWorld Agent · Qwen 3.7-Plus · Thinking · StandardAlibaba Qwen · Qwen 3.7-Plus#10Binary completion: 2.8%OSWorld 2.0 computer-use agent