Agent configuration · SWE-bench Verified

ACoder

The ranked object: model, scaffold, tools, permissions, budget, and environment together.

System identity

Agent systemACoder
Base modelClaude 4 Sonnet + Claude 4.1 Opus + GPT-5 + Gemini 2.5 Pro
Source model ID or labelclaude-4-sonnet; claude-4.1-opus; gpt-5-0807-global; gemini-2.5-pro-06-17
ScaffoldACoder
Scaffold versionNot reported
ToolsNot reported
Comparability groupswe-bench-verified-full-500

Evaluation configuration

Source snapshotSWE-bench Verified · Verified · Full leaderboard · 500 tasks · Jul 30, 2026
Tasks500 human-validated GitHub issues
EnvironmentReproducible repository environments evaluated by SWE-bench
BudgetLimits vary by submitted configuration
Evaluation date
Data captured
Resolved76.4%
Partial scoreNot reported

Published evidence

Open source
EvidenceListed by the benchmark publisher
SourceSWE-bench Verified
CostNot reported
Cost basisThe source does not report a comparable cost for this configuration

Comparable cost is not reported for this configuration, so it is not treated as a low-cost option.

What this result does not prove

This result applies to the exact system and source conditions shown here. It does not establish a universal best agent, production reliability, data governance, or performance on a different benchmark version.

All AI Agent rankings