Model rankings / AI Agent leaderboard PUBLISHED AGENT EVALUATIONS
AI Agent leaderboard 66 Agent systems74 Agent benchmark configurations109 BFCL tool-calling model results8 Benchmark scope
</> Coding agents 21 >_ Terminal & DevOps 7 ◎ Browser agents 7 ⌘ Computer-use agents 6 { } Tools & APIs 7 ⌕ Research agents 9 ↻ Business workflows 9
Directory inclusion verifies identity and an official entry point—not performance or a recommendation. Only same-scope public evidence appears in the ranking below.
Atlassian Rovo Dev Atlassian
Agent product Commercial
Sources checked: Aug 2, 2026
Augment Agent Augment Code
Agent product Commercial
Source reachable
Sources checked: Aug 2, 2026
Claude Code Anthropic
Agent product Commercial
Sources checked: Aug 2, 2026
Cline Cline Bot
Open-source Agent
Source reachable
Sources checked: Aug 2, 2026
Codebuff CodebuffAI
Open-source Agent
Source reachable
Sources checked: Aug 2, 2026
Cursor Agent Anysphere
Agent product Commercial
Source reachable
Sources checked: Aug 2, 2026
Devin Cognition
Agent product Commercial
Source reachable
Sources checked: Aug 2, 2026
Factory Droid Factory
Agent product Commercial
Source reachable
Sources checked: Aug 2, 2026
Show all 21 Saved on this device
Coding agents 10 comparable configurations
Published results SWE-bench Verified Verified · Full leaderboard · 500 tasks Historical evidence Source data date Not published by source Retrieved by OpenGPT Jul 30, 2026 Ranking last changed No comparable history yet OpenGPT checks this Agent source at least once a week. Source result Resolved
Tasks 500 · human-validated GitHub issues
Evidence level External caution
Open source↗ Evidence note Keep these results as historical evidence. A 2026 OpenAI audit reported design and contamination concerns in SWE-bench Verified.
Open audit↗ 1
Claude 4.5 Opus medium (20251101) · live-SWE-agent Listed by the benchmark publisher Resolved 79.2%
Repeat-run evidence Not published
Reported benchmark cost Not published Cost per successful task: Not published
1
Claude 4.5 Opus · Sonar Foundation Agent Listed by the benchmark publisher Resolved 79.2%
Repeat-run evidence Not published
Reported benchmark cost Not published Cost per successful task: Not published
3
Doubao-Seed-Code + Doubao-Seed-1.6 · TRAE Listed by the benchmark publisher Resolved 78.8%
Repeat-run evidence Not published
Reported benchmark cost Not published Cost per successful task: Not published
4
Gemini 3 Pro Preview (2025-11-18) · live-SWE-agent Listed by the benchmark publisher Resolved 77.4%
Repeat-run evidence Not published
Reported benchmark cost Not published Cost per successful task: Not published
5
Claude Sonnet 4 + GPT-5 · Atlassian Rovo Dev · 2025-09-02 Listed by the benchmark publisher Resolved 76.8%
Repeat-run evidence Not published
Reported benchmark cost Not published Cost per successful task: Not published
5
Claude 4 Sonnet · EPAM AI/Run Developer Agent · v20250719 Listed by the benchmark publisher Resolved 76.8%
Repeat-run evidence Not published
Reported benchmark cost Not published Cost per successful task: Not published
5
Claude 4.5 Opus (high reasoning) · mini-SWE-agent · 2.0.0 Listed by the benchmark publisher Resolved 76.8%
Repeat-run evidence Not published
Reported benchmark cost $376.95 Cost per successful task: Not published
8
Claude 4 Sonnet + Claude 4.1 Opus + GPT-5 + Gemini 2.5 Pro · ACoder Listed by the benchmark publisher Resolved 76.4%
Repeat-run evidence Not published
Reported benchmark cost Not published Cost per successful task: Not published
9
Gemini 3 Flash (high reasoning) · mini-SWE-agent · 2.0.0 Listed by the benchmark publisher Resolved 75.8%
Repeat-run evidence Not published
Reported benchmark cost $177.98 Cost per successful task: Not published
9
MiniMax M2.5 (high reasoning) · mini-SWE-agent · 2.0.0 Listed by the benchmark publisher Resolved 75.8%
Repeat-run evidence Not published
Reported benchmark cost $36.64 Cost per successful task: Not published
Saved on this device
NO-API STARTING POINT
Check adoption risks These choices reorder published configurations and flag adoption risks. They do not manufacture missing evidence.
Share selection criteria↗ Deployment risk check Flexible Prefer self-managed Hosted is acceptable Data-handling risk check Standard data Sensitive or regulated data Priority Task completion
Best starting points for these priorities 3 comparable configurations
live-SWE-agent + Claude 4.5 Opus medium (20251101) 79.2% · Not published cost per successful task Higher source result under this benchmark configuration.
Agent details → Sonar Foundation Agent + Claude 4.5 Opus 79.2% · Not published cost per successful task Higher source result under this benchmark configuration.
Agent details → TRAE + Doubao-Seed-Code 78.8% · Not published cost per successful task Higher source result under this benchmark configuration.
Agent details → Source-reported data; missing values remain unknown.
Validate before adoption Public rankings create a shortlist. A small private trial decides whether the configuration fits your work.
1 Choose 10–20 representative tasks and define a clear pass condition.
2 Lock the agent, model, tools, permissions, and budget before testing.
3 Run each important task more than once; record success, cost, time, and takeovers.
4 Review failures and data-handling risks before a production decision.