The ranked object: model, scaffold, tools, permissions, budget, and environment together.
PUBLISHED AGENT EVALUATIONS
AI Agent leaderboard
Compare reported agent configurations within the same published benchmark scope. Each result keeps the original scale and every configuration field made public by the source.
There is no cross-benchmark overall score. OpenGPT organizes upstream results; it did not rerun these evaluations.
One component of the system. An agent result is not attributed to the model alone.
NO-API STARTING POINT
Find agent configurations to investigate
Task and priority order the candidates. Deployment and data-boundary answers add review warnings. Only configurations in the published snapshots are returned.
Choose a task category
Coding agents
SWE-bench Verified
Verified · Full leaderboard · 500 tasks · Data captured Jul 30, 2026
Source resultResolved
Tasks500 human-validated GitHub issues
EnvironmentReproducible repository environments evaluated by SWE-bench
BudgetLimits vary by submitted configuration
1
live-SWE-agent + Claude 4.5 Opus medium (20251101)Base model: Claude 4.5 Opus medium (20251101)Source model ID or label: claude-opus-4-5-20251101Listed by the benchmark publisher
Scaffold: live-SWE-agentScaffold version: Not reportedTools: Not reportedVerified · Full leaderboard · 500 tasks
Resolved79.2%
CostNot reportedDec 15, 2025
1
Sonar Foundation Agent + Claude 4.5 OpusBase model: Claude 4.5 OpusSource model ID or label: claude-opus-4-5Listed by the benchmark publisher
Scaffold: Sonar Foundation AgentScaffold version: Not reportedTools: Not reportedVerified · Full leaderboard · 500 tasks
Resolved79.2%
CostNot reportedDec 5, 2025
3
TRAE + Doubao-Seed-CodeBase model: Doubao-Seed-Code + Doubao-Seed-1.6Source model ID or label: Doubao-Seed-Code; Doubao-Seed-1.6Listed by the benchmark publisher
Scaffold: TRAEScaffold version: Not reportedTools: Not reportedVerified · Full leaderboard · 500 tasks
Resolved78.8%
CostNot reportedSep 28, 2025
4
live-SWE-agent + Gemini 3 Pro Preview (2025-11-18)Base model: Gemini 3 Pro Preview (2025-11-18)Source model ID or label: gemini-3-pro-previewListed by the benchmark publisher
Scaffold: live-SWE-agentScaffold version: Not reportedTools: Not reportedVerified · Full leaderboard · 500 tasks
Resolved77.4%
CostNot reportedNov 20, 2025
5
Atlassian Rovo Dev (2025-09-02)Base model: Claude Sonnet 4 + GPT-5Source model ID or label: claude-sonnet-4-20250514; gpt-5Listed by the benchmark publisher
Scaffold: Atlassian Rovo DevScaffold version: 2025-09-02Tools: Not reportedVerified · Full leaderboard · 500 tasks
Resolved76.8%
CostNot reportedSep 2, 2025
5
EPAM AI/Run Developer Agent v20250719 + Claude 4 SonnetBase model: Claude 4 SonnetSource model ID or label: claude-sonnet-4-20250514Listed by the benchmark publisher
Scaffold: EPAM AI/Run Developer AgentScaffold version: v20250719Tools: Not reportedVerified · Full leaderboard · 500 tasks
Resolved76.8%
CostNot reportedAug 4, 2025
5
mini-SWE-agent + Claude 4.5 Opus (high reasoning)Base model: Claude 4.5 Opus (high reasoning)Source model ID or label: claude-4-5-opusListed by the benchmark publisher
Scaffold: mini-SWE-agentScaffold version: 2.0.0Tools: Minimal repository shellVerified · Full leaderboard · 500 tasks
Resolved76.8%
Cost$376.95Feb 17, 2026
8
ACoderBase model: Claude 4 Sonnet + Claude 4.1 Opus + GPT-5 + Gemini 2.5 ProSource model ID or label: claude-4-sonnet; claude-4.1-opus; gpt-5-0807-global; gemini-2.5-pro-06-17Listed by the benchmark publisher
Scaffold: ACoderScaffold version: Not reportedTools: Not reportedVerified · Full leaderboard · 500 tasks
Resolved76.4%
CostNot reportedAug 19, 2025
9
mini-SWE-agent + Gemini 3 Flash (high reasoning)Base model: Gemini 3 Flash (high reasoning)Source model ID or label: gemini-3-flash-previewListed by the benchmark publisher
Scaffold: mini-SWE-agentScaffold version: 2.0.0Tools: Minimal repository shellVerified · Full leaderboard · 500 tasks
Resolved75.8%
Cost$177.98Feb 17, 2026
9
mini-SWE-agent + MiniMax M2.5 (high reasoning)Base model: MiniMax M2.5 (high reasoning)Source model ID or label: minimax-m2.5Listed by the benchmark publisher
Scaffold: mini-SWE-agentScaffold version: 2.0.0Tools: Minimal repository shellVerified · Full leaderboard · 500 tasks
Resolved75.8%
Cost$36.64Feb 17, 2026