Agent 評測
OSWorld 2.0 Agent 評測
OpenGPT 將10筆公開記錄保留在下方所示的 OSWorld 2.0 評測版本與環境內。
- 工作
- 108
- 公開記錄
- 10
- 來源資料日期
- 身分核驗日期
- 目前
資料範圍
task version v2026.06.24 · 500-step budget. long-horizon computer workflows. Desktop applications and operating-system workflows in reproducible VMs. 500 steps.
OpenGPT 如何處理這些資料
OpenGPT 保留發布方的指標、工作範圍、環境、版本、設定欄位與資料日期,不產生跨評測 Agent 總分。
可比範圍
- 更高結果只適用於所展示的評測版本與條件。
- Agent 結果反映完整系統,而不是基礎模型單獨能力。
- 成本、穩定性、安全性或速度沒有公開時保持未報告,不進行推測。
- Rows preserve the complete agent, model, tool mode, task version, step budget, and source-reported metrics.
task version v2026.06.24 · 500-step budget
展示10筆公開記錄中的10筆;每項結果都保留來源評測與設定。
- OSWorld Agent · Claude Opus 4.8 · Max · Batched toolAnthropic · Claude Opus 4.8#1Binary completion: 20.6%OSWorld 2.0 computer-use agent
- OSWorld Agent · Claude Opus 4.8 · Max · StandardAnthropic · Claude Opus 4.8#2Binary completion: 18.52%OSWorld 2.0 computer-use agent
- OSWorld Agent · Claude Opus 4.7 · Max · Batched toolAnthropic · Claude Opus 4.7#3Binary completion: 18.2%OSWorld 2.0 computer-use agent
- OSWorld Agent · Claude Opus 4.7 · Max · StandardAnthropic · Claude Opus 4.7#4Binary completion: 13.9%OSWorld 2.0 computer-use agent
- OSWorld Agent · GPT-5.5 · Xhigh · Batch toolOpenAI · GPT-5.5#5Binary completion: 13.0%OSWorld 2.0 computer-use agent
- OSWorld Agent · Claude Sonnet 4.6 · Medium · StandardAnthropic · Claude Sonnet 4.6#6Binary completion: 9.3%OSWorld 2.0 computer-use agent
- OSWorld Agent · Claude Sonnet 4.6 · Max · StandardAnthropic · Claude Sonnet 4.6#7Binary completion: 8.3%OSWorld 2.0 computer-use agent
- OSWorld Agent · MiniMax M3 · Enabled · StandardMiniMax · MiniMax M3#8Binary completion: 4.6%OSWorld 2.0 computer-use agent
- OSWorld Agent · Kimi 2.6 · Enabled · StandardMoonshot AI · Kimi 2.6#8Binary completion: 4.6%OSWorld 2.0 computer-use agent
- OSWorld Agent · Qwen 3.7-Plus · Thinking · StandardAlibaba Qwen · Qwen 3.7-Plus#10Binary completion: 2.8%OSWorld 2.0 computer-use agent