Shared benchmark scope · SWE-bench Verified

Same-benchmark agent comparison

The entries below share the same benchmark scope. Their exact scaffolds, model versions, dates, and original source values remain visible.

A higher source value only means better performance on this benchmark and configuration. It is not a cross-benchmark or production-readiness score.

Compare like with like

Decision overview

Only configurations from the same benchmark version and task scope can be placed in one comparison.

Source-reported data; missing values remain unknown.
SWE-agentmini-SWE-agent + MiniMax M2.5 (high reasoning)
Resolved75.8%Cost per successful taskNot publishedPublished result rangeNot published
Agent details
SWE-agentmini-SWE-agent + Claude 4.5 Opus (high reasoning)
Resolved76.8%Cost per successful taskNot publishedPublished result rangeNot published
Agent details
17 identical or entirely unreported fields hidden
Agent systemmini-SWE-agent + MiniMax M2.5 (high reasoning)mini-SWE-agent + Claude 4.5 Opus (high reasoning)
Configuration
Base modelMiniMax M2.5 (high reasoning)Claude 4.5 Opus (high reasoning)
Source model ID or labelminimax-m2.5claude-4-5-opus
Outcome
Resolved75.8%76.8%
Reported benchmark cost$36.64$376.95
Source
Open sourceOpen sourceOpen source

Configuration

mini-SWE-agent + MiniMax M2.5 (high reasoning)

Base model
MiniMax M2.5 (high reasoning)
Source model ID or label
minimax-m2.5

mini-SWE-agent + Claude 4.5 Opus (high reasoning)

Base model
Claude 4.5 Opus (high reasoning)
Source model ID or label
claude-4-5-opus

Outcome

mini-SWE-agent + MiniMax M2.5 (high reasoning)

Resolved
75.8%
Reported benchmark cost
$36.64

mini-SWE-agent + Claude 4.5 Opus (high reasoning)

Resolved
76.8%
Reported benchmark cost
$376.95

Source

Source-reported data; missing values remain unknown.

Validate before adoption

Public rankings create a shortlist. A small private trial decides whether the configuration fits your work.

  1. 1

    Choose 10–20 representative tasks and define a clear pass condition.

  2. 2

    Lock the agent, model, tools, permissions, and budget before testing.

  3. 3

    Run each important task more than once; record success, cost, time, and takeovers.

  4. 4

    Review failures and data-handling risks before a production decision.

Agent directory and evaluations