에이전트 벤치마크

SWE-bench Verified 에이전트 벤치마크

SWE-bench Verified의 표시된 버전과 환경에 연결된 공개 행 10개를 유지합니다.

과제
500
공개 행
10
출처 데이터 날짜
ID 검토일
과거 기록

근거 범위

Verified · Top 10 published configurations · 500 tasks. human-validated GitHub issues. Reproducible repository environments evaluated by the upstream benchmark. Configuration-specific upstream limits.

OpenGPT의 근거 처리 방식

게시자의 지표, 과제 범위, 환경, 버전, 구성과 날짜를 보존하며 벤치마크 간 종합 점수를 만들지 않습니다.

비교 범위

  • 높은 결과는 표시된 벤치마크 버전과 조건에만 적용됩니다.
  • 에이전트 결과는 기반 모델 하나가 아닌 전체 시스템을 반영합니다.
  • 비용, 신뢰성, 안전성 또는 지연 시간 누락값은 추정하지 않습니다.
  • Keep these published results as historical evidence, not a current universal measure of coding ability. A 2026 OpenAI audit reported fundamental design and contamination concerns in SWE-bench Verified.

Verified · Top 10 published configurations · 500 tasks

공개 10개 중 10개를 표시합니다. 모든 결과는 출처 벤치마크와 구성에 연결됩니다.

  1. live-SWE-agent + Claude 4.5 Opus medium (20251101)UIUC · Claude 4.5 Opus medium (20251101)#1Resolved: 79.2%live-SWE-agent
  2. Sonar Foundation Agent + Claude 4.5 OpusSonar · Claude 4.5 Opus#1Resolved: 79.2%Sonar Foundation Agent
  3. TRAE + Doubao-Seed-CodeByteDance · Doubao-Seed-Code + Doubao-Seed-1.6#3Resolved: 78.8%TRAE
  4. live-SWE-agent + Gemini 3 Pro Preview (2025-11-18)UIUC · Gemini 3 Pro Preview (2025-11-18)#4Resolved: 77.4%live-SWE-agent
  5. Atlassian Rovo Dev (2025-09-02)Atlassian · Claude Sonnet 4 + GPT-5#5Resolved: 76.8%Atlassian Rovo Dev
  6. EPAM AI/Run Developer Agent v20250719 + Claude 4 SonnetEPAM Systems · Claude 4 Sonnet#5Resolved: 76.8%EPAM AI/Run Developer Agent
  7. mini-SWE-agent + Claude 4.5 Opus (high reasoning)SWE-agent · Claude 4.5 Opus (high reasoning)#5Resolved: 76.8%mini-SWE-agent
  8. ACoderACoder · Claude 4 Sonnet + Claude 4.1 Opus + GPT-5 + Gemini 2.5 Pro#8Resolved: 76.4%ACoder
  9. mini-SWE-agent + Gemini 3 Flash (high reasoning)SWE-agent · Gemini 3 Flash (high reasoning)#9Resolved: 75.8%mini-SWE-agent
  10. mini-SWE-agent + MiniMax M2.5 (high reasoning)SWE-agent · MiniMax M2.5 (high reasoning)#9Resolved: 75.8%mini-SWE-agent