共同评测范围 · OSWorld 2.0

同评测Agent对比

以下配置来自同一个评测范围,并保留各自的框架、模型版本、日期和来源原始值。

来源值较高只代表这个评测和配置下的结果,不是跨评测总分,也不代表已经适合生产使用。

只比较相同范围

决策概览

只有评测版本和任务范围一致的配置,才能加入同一组对比。

数据来自公开来源,缺失项保持未知。
Moonshot AIOSWorld Agent · Kimi 2.6 · Enabled · Standard
完全完成率4.6%每次成功成本未公布来源公布的结果范围未公布
Agent详情
AnthropicOSWorld Agent · Claude Opus 4.8 · Max · Batched tool
完全完成率20.6%每次成功成本未公布来源公布的结果范围未公布
Agent详情
已隐藏15项完全相同或全部未提供的字段
Agent系统OSWorld Agent · Kimi 2.6 · Enabled · StandardOSWorld Agent · Claude Opus 4.8 · Max · Batched tool
配置
基础模型Kimi 2.6Claude Opus 4.8
来源模型ID或名称Kimi 2.6Claude Opus 4.8
工具Standard computer-use toolBatched computer-use tool
结果
完全完成率4.6%20.6%
部分完成分22.1%54.8%
来源公布的评测总成本US$708未公布
数据来源
打开来源打开来源打开来源

配置

OSWorld Agent · Kimi 2.6 · Enabled · Standard

基础模型
Kimi 2.6
来源模型ID或名称
Kimi 2.6
工具
Standard computer-use tool

OSWorld Agent · Claude Opus 4.8 · Max · Batched tool

基础模型
Claude Opus 4.8
来源模型ID或名称
Claude Opus 4.8
工具
Batched computer-use tool

结果

OSWorld Agent · Kimi 2.6 · Enabled · Standard

完全完成率
4.6%
部分完成分
22.1%
来源公布的评测总成本
US$708

OSWorld Agent · Claude Opus 4.8 · Max · Batched tool

完全完成率
20.6%
部分完成分
54.8%
来源公布的评测总成本
未公布

数据来源

数据来自公开来源,缺失项保持未知。

采用前做一次真实验证

公开榜单用于缩小候选范围,是否适合你的工作,需要通过小规模私有测试确认。

  1. 1

    选择 10–20 个代表性任务,并写清通过条件。

  2. 2

    测试前锁定Agent、模型、工具、权限和预算。

  3. 3

    重要任务重复运行,记录成功率、成本、时间和人工接管。

  4. 4

    检查失败案例和数据风险后,再决定是否正式采用。

Agent目录与评测