Shared benchmark scope · SWE-bench Verified

Same-benchmark agent comparison

The entries below share the same benchmark scope. Their exact scaffolds, model versions, dates, and original source values remain visible.

A higher source value only means better performance on this benchmark and configuration. It is not a cross-benchmark or production-readiness score.

Compare like with like

Decision overview

Only configurations from the same benchmark version and task scope can be placed in one comparison.

Source-reported data; missing values remain unknown.
UIUClive-SWE-agent + Claude 4.5 Opus medium (20251101)
Resolved79.2%Cost per successful taskNot publishedPublished result rangeNot published
Agent details
UIUClive-SWE-agent + Gemini 3 Pro Preview (2025-11-18)
Resolved77.4%Cost per successful taskNot publishedPublished result rangeNot published
Agent details
17 identical or entirely unreported fields hidden
Agent systemlive-SWE-agent + Claude 4.5 Opus medium (20251101)live-SWE-agent + Gemini 3 Pro Preview (2025-11-18)
Configuration
Base modelClaude 4.5 Opus medium (20251101)Gemini 3 Pro Preview (2025-11-18)
Source model ID or labelclaude-opus-4-5-20251101gemini-3-pro-preview
Outcome
Resolved79.2%77.4%
Trust
Evaluation dateDec 15, 2025Nov 20, 2025
Source
Open sourceOpen sourceOpen source

Configuration

live-SWE-agent + Claude 4.5 Opus medium (20251101)

Base model
Claude 4.5 Opus medium (20251101)
Source model ID or label
claude-opus-4-5-20251101

live-SWE-agent + Gemini 3 Pro Preview (2025-11-18)

Base model
Gemini 3 Pro Preview (2025-11-18)
Source model ID or label
gemini-3-pro-preview

Outcome

live-SWE-agent + Claude 4.5 Opus medium (20251101)

Resolved
79.2%

live-SWE-agent + Gemini 3 Pro Preview (2025-11-18)

Resolved
77.4%

Trust

live-SWE-agent + Claude 4.5 Opus medium (20251101)

Evaluation date
Dec 15, 2025

live-SWE-agent + Gemini 3 Pro Preview (2025-11-18)

Evaluation date
Nov 20, 2025

Source

Source-reported data; missing values remain unknown.

Validate before adoption

Public rankings create a shortlist. A small private trial decides whether the configuration fits your work.

  1. 1

    Choose 10–20 representative tasks and define a clear pass condition.

  2. 2

    Lock the agent, model, tools, permissions, and budget before testing.

  3. 3

    Run each important task more than once; record success, cost, time, and takeovers.

  4. 4

    Review failures and data-handling risks before a production decision.

Agent directory and evaluations