Shared benchmark scope · SWE-bench Verified

Same-benchmark agent comparison

The entries below share the same benchmark scope. Their exact scaffolds, model versions, dates, and original source values remain visible.

A higher source value only means better performance on this benchmark and configuration. It is not a cross-benchmark or production-readiness score.

Compare like with like

Decision overview

Only configurations from the same benchmark version and task scope can be placed in one comparison.

Source-reported data; missing values remain unknown.
UIUClive-SWE-agent + Gemini 3 Pro Preview (2025-11-18)
Resolved77.4%Cost per successful taskNot publishedPublished result rangeNot published
Agent details
UIUClive-SWE-agent + Claude 4.5 Opus medium (20251101)
Resolved79.2%Cost per successful taskNot publishedPublished result rangeNot published
Agent details
17 identical or entirely unreported fields hidden
Agent systemlive-SWE-agent + Gemini 3 Pro Preview (2025-11-18)live-SWE-agent + Claude 4.5 Opus medium (20251101)
Configuration
Base modelGemini 3 Pro Preview (2025-11-18)Claude 4.5 Opus medium (20251101)
Source model ID or labelgemini-3-pro-previewclaude-opus-4-5-20251101
Outcome
Resolved77.4%79.2%
Trust
Evaluation dateNov 20, 2025Dec 15, 2025
Source
Open sourceOpen sourceOpen source

Configuration

live-SWE-agent + Gemini 3 Pro Preview (2025-11-18)

Base model
Gemini 3 Pro Preview (2025-11-18)
Source model ID or label
gemini-3-pro-preview

live-SWE-agent + Claude 4.5 Opus medium (20251101)

Base model
Claude 4.5 Opus medium (20251101)
Source model ID or label
claude-opus-4-5-20251101

Outcome

live-SWE-agent + Gemini 3 Pro Preview (2025-11-18)

Resolved
77.4%

live-SWE-agent + Claude 4.5 Opus medium (20251101)

Resolved
79.2%

Trust

live-SWE-agent + Gemini 3 Pro Preview (2025-11-18)

Evaluation date
Nov 20, 2025

live-SWE-agent + Claude 4.5 Opus medium (20251101)

Evaluation date
Dec 15, 2025

Source

Source-reported data; missing values remain unknown.

Validate before adoption

Public rankings create a shortlist. A small private trial decides whether the configuration fits your work.

  1. 1

    Choose 10–20 representative tasks and define a clear pass condition.

  2. 2

    Lock the agent, model, tools, permissions, and budget before testing.

  3. 3

    Run each important task more than once; record success, cost, time, and takeovers.

  4. 4

    Review failures and data-handling risks before a production decision.

Agent directory and evaluations