Shared benchmark scope · τ²-bench Core

Same-benchmark agent comparison

The entries below share the same benchmark scope. Their exact scaffolds, model versions, dates, and original source values remain visible.

A higher source value only means better performance on this benchmark and configuration. It is not a cross-benchmark or production-readiness score.

Compare like with like

Decision overview

Only configurations from the same benchmark version and task scope can be placed in one comparison.

Source-reported data; missing values remain unknown.
Zhipu AIτ² standard agent · GLM-5 · thinking
Average Pass¹81.01%Cost per successful taskNot publishedPublished result rangeNot published
Agent details
Alibaba Cloudτ² standard agent · Qwen3.5-397B-A17B · thinking
Average Pass¹87.91%Cost per successful taskNot publishedPublished result rangeNot published
Agent details
18 identical or entirely unreported fields hidden
Agent systemτ² standard agent · GLM-5 · thinkingτ² standard agent · Qwen3.5-397B-A17B · thinking
Configuration
Base modelGLM-5Qwen3.5-397B-A17B
Source model ID or labelGLM-5 · thinkingQwen3.5-397B-A17B · thinking
Outcome
Average Pass¹81.01%87.91%
Source
Open sourceOpen sourceOpen source

Configuration

τ² standard agent · GLM-5 · thinking

Base model
GLM-5
Source model ID or label
GLM-5 · thinking

τ² standard agent · Qwen3.5-397B-A17B · thinking

Base model
Qwen3.5-397B-A17B
Source model ID or label
Qwen3.5-397B-A17B · thinking

Outcome

τ² standard agent · GLM-5 · thinking

Average Pass¹
81.01%

τ² standard agent · Qwen3.5-397B-A17B · thinking

Average Pass¹
87.91%

Source

Source-reported data; missing values remain unknown.

Validate before adoption

Public rankings create a shortlist. A small private trial decides whether the configuration fits your work.

  1. 1

    Choose 10–20 representative tasks and define a clear pass condition.

  2. 2

    Lock the agent, model, tools, permissions, and budget before testing.

  3. 3

    Run each important task more than once; record success, cost, time, and takeovers.

  4. 4

    Review failures and data-handling risks before a production decision.

Agent directory and evaluations