에이전트 벤치마크
OSWorld 2.0 에이전트 벤치마크
OSWorld 2.0의 표시된 버전과 환경에 연결된 공개 행 10개를 유지합니다.
- 과제
- 108
- 공개 행
- 10
- 출처 데이터 날짜
- ID 검토일
- 현재
근거 범위
task version v2026.06.24 · 500-step budget. long-horizon computer workflows. Desktop applications and operating-system workflows in reproducible VMs. 500 steps.
OpenGPT의 근거 처리 방식
게시자의 지표, 과제 범위, 환경, 버전, 구성과 날짜를 보존하며 벤치마크 간 종합 점수를 만들지 않습니다.
비교 범위
- 높은 결과는 표시된 벤치마크 버전과 조건에만 적용됩니다.
- 에이전트 결과는 기반 모델 하나가 아닌 전체 시스템을 반영합니다.
- 비용, 신뢰성, 안전성 또는 지연 시간 누락값은 추정하지 않습니다.
- Rows preserve the complete agent, model, tool mode, task version, step budget, and source-reported metrics.
task version v2026.06.24 · 500-step budget
공개 10개 중 10개를 표시합니다. 모든 결과는 출처 벤치마크와 구성에 연결됩니다.
- OSWorld Agent · Claude Opus 4.8 · Max · Batched toolAnthropic · Claude Opus 4.8#1Binary completion: 20.6%OSWorld 2.0 computer-use agent
- OSWorld Agent · Claude Opus 4.8 · Max · StandardAnthropic · Claude Opus 4.8#2Binary completion: 18.52%OSWorld 2.0 computer-use agent
- OSWorld Agent · Claude Opus 4.7 · Max · Batched toolAnthropic · Claude Opus 4.7#3Binary completion: 18.2%OSWorld 2.0 computer-use agent
- OSWorld Agent · Claude Opus 4.7 · Max · StandardAnthropic · Claude Opus 4.7#4Binary completion: 13.9%OSWorld 2.0 computer-use agent
- OSWorld Agent · GPT-5.5 · Xhigh · Batch toolOpenAI · GPT-5.5#5Binary completion: 13.0%OSWorld 2.0 computer-use agent
- OSWorld Agent · Claude Sonnet 4.6 · Medium · StandardAnthropic · Claude Sonnet 4.6#6Binary completion: 9.3%OSWorld 2.0 computer-use agent
- OSWorld Agent · Claude Sonnet 4.6 · Max · StandardAnthropic · Claude Sonnet 4.6#7Binary completion: 8.3%OSWorld 2.0 computer-use agent
- OSWorld Agent · MiniMax M3 · Enabled · StandardMiniMax · MiniMax M3#8Binary completion: 4.6%OSWorld 2.0 computer-use agent
- OSWorld Agent · Kimi 2.6 · Enabled · StandardMoonshot AI · Kimi 2.6#8Binary completion: 4.6%OSWorld 2.0 computer-use agent
- OSWorld Agent · Qwen 3.7-Plus · Thinking · StandardAlibaba Qwen · Qwen 3.7-Plus#10Binary completion: 2.8%OSWorld 2.0 computer-use agent