Agent configuration · SWE-bench Verified

live-SWE-agent + Claude 4.5 Opus medium (20251101)

The ranked object: model, scaffold, tools, permissions, budget, and environment together.

Resolved79.2%
Cost per successful taskNot published
Published result rangeNot published

01

Configuration

Agent systemlive-SWE-agent + Claude 4.5 Opus medium (20251101)
Base modelClaude 4.5 Opus medium (20251101)
Source model ID or labelclaude-opus-4-5-20251101
Scaffoldlive-SWE-agent
Scaffold versionNot reported
ToolsNot reported
EnvironmentReproducible repository environments evaluated by SWE-bench
BudgetLimits vary by submitted configuration

02

Outcome

Resolved79.2%
Partial scoreNot published
Cost per successful taskNot published
Reported benchmark costNot reported
Median completion timeNot published
Human interventionNot published

03

Operation

Repeat-run evidenceNot published
Published result rangeNot published
Safety noteNot published
Comparability groupswe-bench-verified-full-500
Evaluation date
Data captured

04

Trust

EvidenceListed by the benchmark publisher
SourceSWE-bench Verified
Source snapshotVerified · Full leaderboard · 500 tasks
Source data date
Retrieved by OpenGPT
Ranking last changedNo comparable history yet
Last checked
Tasks500 human-validated GitHub issues
Evidence levelExternal caution
Evidence note

Keep these results as historical evidence. A 2026 OpenAI audit reported design and contamination concerns in SWE-bench Verified.

Open audit

What this result does not prove

This result applies to the exact system and source conditions shown here. It does not establish a universal best agent, production reliability, data governance, or performance on a different benchmark version.

Compare like with like

Only configurations from the same benchmark version and task scope can be placed in one comparison.

Source-reported data; missing values remain unknown.

Validate before adoption

Public rankings create a shortlist. A small private trial decides whether the configuration fits your work.

  1. 1

    Choose 10–20 representative tasks and define a clear pass condition.

  2. 2

    Lock the agent, model, tools, permissions, and budget before testing.

  3. 3

    Run each important task more than once; record success, cost, time, and takeovers.

  4. 4

    Review failures and data-handling risks before a production decision.

Agent directory and evaluations