Agent configuration · τ²-bench Core

τ² standard agent · Qwen3.5-397B-A17B · thinking

The ranked object: model, scaffold, tools, permissions, budget, and environment together.

Average Pass¹87.91%
Cost per successful taskNot published
Published result rangeNot published

01

Configuration

Agent systemτ² standard agent · Qwen3.5-397B-A17B · thinking
Base modelQwen3.5-397B-A17B
Source model ID or labelQwen3.5-397B-A17B · thinking
Scaffoldτ² standard agent
Scaffold versionv1.0.1
ToolsStandard retail, airline, and telecom domain tools
EnvironmentStandard agent and domain tools with a shared GPT-5.2 low-reasoning user simulator
Budget4 trials per task · 1,112 trajectories per configuration

02

Outcome

Average Pass¹87.91%
Partial scoreNot published
Cost per successful taskNot published
Reported benchmark costNot reported
Median completion timeNot published
Human interventionNot published

03

Operation

Repeat-run evidence4 runs
Published result rangeNot published
Safety noteNot published
Comparability grouptau2-v1-0-1-base-standard-gpt-5-2-low-4-trials
Evaluation date
Data captured

04

Trust

EvidenceListed by the benchmark publisher
SourceSierra Research · τ²-bench
Source snapshotv1.0.1 · standard agent · GPT-5.2 low user simulator · 4 trials
Source data dateNot published by source
Retrieved by OpenGPT
Ranking last changedNo comparable history yet
Last checked
Tasks278 base tasks across retail, airline, and telecom
Evidence levelPublisher maintained
Evidence note

Only eight publisher-submitted rows with the same v1.0.1 benchmark, standard scaffold and tools, user simulator, base task splits, and four-trial protocol are included. Other τ-bench rows stay outside this scope.

Open audit

What this result does not prove

This result applies to the exact system and source conditions shown here. It does not establish a universal best agent, production reliability, data governance, or performance on a different benchmark version.

Compare like with like

Only configurations from the same benchmark version and task scope can be placed in one comparison.

Source-reported data; missing values remain unknown.

Validate before adoption

Public rankings create a shortlist. A small private trial decides whether the configuration fits your work.

  1. 1

    Choose 10–20 representative tasks and define a clear pass condition.

  2. 2

    Lock the agent, model, tools, permissions, and budget before testing.

  3. 3

    Run each important task more than once; record success, cost, time, and takeovers.

  4. 4

    Review failures and data-handling risks before a production decision.

Agent directory and evaluations