PUBLISHED AGENT EVALUATIONS

AI Agent leaderboard

66 Agent systems74 Agent benchmark configurations109 BFCL tool-calling model results8 Benchmark scope

Agent directory

21

Distinct Agent systems with a verified first-party product page, documentation page, or repository.

Directory inclusion verifies identity and an official entry point—not performance or a recommendation. Only same-scope public evidence appears in the ranking below.

Augment AgentAugment Code
Agent productCommercial
Source reachable

Sources checked:

Product details
ClineCline Bot
Open-source Agent
Source reachable

Sources checked:

Product details
CodebuffCodebuffAI
Open-source Agent
Source reachable

Sources checked:

Product details
Cursor AgentAnysphere
Agent productCommercial
Source reachable

Sources checked:

Product details
DevinCognition
Agent productCommercial
Source reachable

Sources checked:

Product details
Factory DroidFactory
Agent productCommercial
Source reachable

Sources checked:

Product details

Saved on this device

Coding agents

10 comparable configurations

Submit Agent
Published resultsSWE-bench VerifiedVerified · Full leaderboard · 500 tasksHistorical evidenceSource data dateNot published by sourceRetrieved by OpenGPTRanking last changedNo comparable history yetOpenGPT checks this Agent source at least once a week.
Source resultResolved
Tasks500 · human-validated GitHub issues
Evidence levelExternal caution
Open source
Evidence note

Keep these results as historical evidence. A 2026 OpenAI audit reported design and contamination concerns in SWE-bench Verified.

Open audit
1
Claude 4.5 Opus · Sonar Foundation AgentListed by the benchmark publisher
Resolved79.2%
Repeat-run evidenceNot published
Reported benchmark costNot publishedCost per successful task: Not published
Details
3
Doubao-Seed-Code + Doubao-Seed-1.6 · TRAEListed by the benchmark publisher
Resolved78.8%
Repeat-run evidenceNot published
Reported benchmark costNot publishedCost per successful task: Not published
Details
5
Claude Sonnet 4 + GPT-5 · Atlassian Rovo Dev · 2025-09-02Listed by the benchmark publisher
Resolved76.8%
Repeat-run evidenceNot published
Reported benchmark costNot publishedCost per successful task: Not published
Details
5
Claude 4.5 Opus (high reasoning) · mini-SWE-agent · 2.0.0Listed by the benchmark publisher
Resolved76.8%
Repeat-run evidenceNot published
Reported benchmark cost$376.95Cost per successful task: Not published
Details
8
ACoderACoder
Claude 4 Sonnet + Claude 4.1 Opus + GPT-5 + Gemini 2.5 Pro · ACoderListed by the benchmark publisher
Resolved76.4%
Repeat-run evidenceNot published
Reported benchmark costNot publishedCost per successful task: Not published
Details
9
Gemini 3 Flash (high reasoning) · mini-SWE-agent · 2.0.0Listed by the benchmark publisher
Resolved75.8%
Repeat-run evidenceNot published
Reported benchmark cost$177.98Cost per successful task: Not published
Details
9
MiniMax M2.5 (high reasoning) · mini-SWE-agent · 2.0.0Listed by the benchmark publisher
Resolved75.8%
Repeat-run evidenceNot published
Reported benchmark cost$36.64Cost per successful task: Not published
Details

Saved on this device

NO-API STARTING POINT

Check adoption risks

These choices reorder published configurations and flag adoption risks. They do not manufacture missing evidence.

Best starting points for these priorities

3 comparable configurations

live-SWE-agent + Claude 4.5 Opus medium (20251101)79.2% · Not published cost per successful task

Higher source result under this benchmark configuration.

Agent details
Sonar Foundation Agent + Claude 4.5 Opus79.2% · Not published cost per successful task

Higher source result under this benchmark configuration.

Agent details
TRAE + Doubao-Seed-Code78.8% · Not published cost per successful task

Higher source result under this benchmark configuration.

Agent details

    Source-reported data; missing values remain unknown.

    Validate before adoption

    Public rankings create a shortlist. A small private trial decides whether the configuration fits your work.

    1. 1

      Choose 10–20 representative tasks and define a clear pass condition.

    2. 2

      Lock the agent, model, tools, permissions, and budget before testing.

    3. 3

      Run each important task more than once; record success, cost, time, and takeovers.

    4. 4

      Review failures and data-handling risks before a production decision.