Agent configuration · τ²-bench Core
τ² standard agent · GPT-5.2 · high
The ranked object: model, scaffold, tools, permissions, budget, and environment together.
01
Configuration
02
Outcome
03
Operation
04
Trust
Evidence note
Only eight publisher-submitted rows with the same v1.0.1 benchmark, standard scaffold and tools, user simulator, base task splits, and four-trial protocol are included. Other τ-bench rows stay outside this scope.
Open audit ↗What this result does not prove
This result applies to the exact system and source conditions shown here. It does not establish a universal best agent, production reliability, data governance, or performance on a different benchmark version.
Source-reported data; missing values remain unknown.
Validate before adoption
Public rankings create a shortlist. A small private trial decides whether the configuration fits your work.
- 1
Choose 10–20 representative tasks and define a clear pass condition.
- 2
Lock the agent, model, tools, permissions, and budget before testing.
- 3
Run each important task more than once; record success, cost, time, and takeovers.
- 4
Review failures and data-handling risks before a production decision.