共同评测范围 · OSWorld 2.0

同评测Agent对比

以下配置来自同一个评测范围,并保留各自的框架、模型版本、日期和来源原始值。

来源值较高只代表这个评测和配置下的结果,不是跨评测总分,也不代表已经适合生产使用。

只比较相同范围

决策概览

只有评测版本和任务范围一致的配置,才能加入同一组对比。

数据来自公开来源,缺失项保持未知。
AnthropicOSWorld Agent · Claude Opus 4.8 · Max · Batched tool
完全完成率20.6%每次成功成本未公布来源公布的结果范围未公布
Agent详情
AnthropicOSWorld Agent · Claude Opus 4.8 · Max · Standard
完全完成率18.52%每次成功成本未公布来源公布的结果范围未公布
Agent详情
已隐藏18项完全相同或全部未提供的字段
Agent系统OSWorld Agent · Claude Opus 4.8 · Max · Batched toolOSWorld Agent · Claude Opus 4.8 · Max · Standard
配置
工具Batched computer-use toolStandard computer-use tool
结果
完全完成率20.6%18.52%
部分完成分54.8%49.33%
数据来源
打开来源打开来源打开来源

配置

OSWorld Agent · Claude Opus 4.8 · Max · Batched tool

工具
Batched computer-use tool

OSWorld Agent · Claude Opus 4.8 · Max · Standard

工具
Standard computer-use tool

结果

OSWorld Agent · Claude Opus 4.8 · Max · Batched tool

完全完成率
20.6%
部分完成分
54.8%

OSWorld Agent · Claude Opus 4.8 · Max · Standard

完全完成率
18.52%
部分完成分
49.33%

数据来源

数据来自公开来源,缺失项保持未知。

采用前做一次真实验证

公开榜单用于缩小候选范围,是否适合你的工作,需要通过小规模私有测试确认。

  1. 1

    选择 10–20 个代表性任务,并写清通过条件。

  2. 2

    测试前锁定Agent、模型、工具、权限和预算。

  3. 3

    重要任务重复运行,记录成功率、成本、时间和人工接管。

  4. 4

    检查失败案例和数据风险后,再决定是否正式采用。

Agent目录与评测