Shared benchmark scope · τ²-bench Core

Same-benchmark agent comparison

The entries below share the same benchmark scope. Their exact scaffolds, model versions, dates, and original source values remain visible.

A higher source value only means better performance on this benchmark and configuration. It is not a cross-benchmark or production-readiness score.

Compare like with like

Decision overview

Only configurations from the same benchmark version and task scope can be placed in one comparison.

Source-reported data; missing values remain unknown.
OpenAIτ² standard agent · GPT-5.2 · high
Average Pass¹84.76%Cost per successful taskNot publishedPublished result rangeNot published
Agent details
Alibaba Cloudτ² standard agent · Qwen3.5-397B-A17B · thinking
Average Pass¹87.91%Cost per successful taskNot publishedPublished result rangeNot published
Agent details
17 identical or entirely unreported fields hidden
Agent systemτ² standard agent · GPT-5.2 · highτ² standard agent · Qwen3.5-397B-A17B · thinking
Configuration
Base modelGPT-5.2Qwen3.5-397B-A17B
Source model ID or labelGPT-5.2 · highQwen3.5-397B-A17B · thinking
Outcome
Average Pass¹84.76%87.91%
Trust
Evaluation dateMay 5, 2026Feb 27, 2026
Source
Open sourceOpen sourceOpen source

Configuration

τ² standard agent · GPT-5.2 · high

Base model
GPT-5.2
Source model ID or label
GPT-5.2 · high

τ² standard agent · Qwen3.5-397B-A17B · thinking

Base model
Qwen3.5-397B-A17B
Source model ID or label
Qwen3.5-397B-A17B · thinking

Outcome

τ² standard agent · GPT-5.2 · high

Average Pass¹
84.76%

τ² standard agent · Qwen3.5-397B-A17B · thinking

Average Pass¹
87.91%

Trust

τ² standard agent · GPT-5.2 · high

Evaluation date
May 5, 2026

τ² standard agent · Qwen3.5-397B-A17B · thinking

Evaluation date
Feb 27, 2026

Source

Source-reported data; missing values remain unknown.

Validate before adoption

Public rankings create a shortlist. A small private trial decides whether the configuration fits your work.

  1. 1

    Choose 10–20 representative tasks and define a clear pass condition.

  2. 2

    Lock the agent, model, tools, permissions, and budget before testing.

  3. 3

    Run each important task more than once; record success, cost, time, and takeovers.

  4. 4

    Review failures and data-handling risks before a production decision.

Agent directory and evaluations