Agent benchmarks
τ²-bench Core agent benchmark
OpenGPT keeps 8 published rows attached to the exact τ²-bench Core benchmark version and environment shown below.
- Tasks
- 278
- Published rows
- 8
- Source data date
- Identity reviewed
- Current
Evidence scope
v1.0.1 · standard agent · GPT-5.2 low user simulator · 4 trials. base tasks across retail, airline, and telecom. Standard agent, standard domain tools, and a shared GPT-5.2 low-reasoning user simulator. 4 trials per task · 1,112 trajectories per configuration.
How OpenGPT handles this evidence
OpenGPT preserves the publisher's metric, task scope, environment, version, configuration fields, and source date. It does not create a cross-benchmark overall Agent score.
Comparison boundaries
- A higher result applies only to the displayed benchmark version and conditions.
- Agent results reflect the full system, not the base model alone.
- Missing cost, reliability, safety, or latency data remains unreported rather than being inferred.
- OpenGPT includes only the eight publisher-submitted rows that share benchmark v1.0.1, the standard scaffold and tools, the same user simulator, the same base task splits, and four trials. Other τ-bench rows remain outside this scope.
v1.0.1 · standard agent · GPT-5.2 low user simulator · 4 trials
Showing 8 of 8 published rows. Every result remains attached to the source benchmark and configuration.
- τ² standard agent · Qwen3.5-397B-A17B · thinkingAlibaba Cloud · Qwen3.5-397B-A17B#1Average Pass¹: 87.91%τ² standard agent
- τ² standard agent · Claude Opus 4.5 · highAnthropic · Claude Opus 4.5#2Average Pass¹: 85.31%τ² standard agent
- τ² standard agent · GPT-5.2 · highOpenAI · GPT-5.2#3Average Pass¹: 84.76%τ² standard agent
- τ² standard agent · Gemini 3 Flash Preview · highGoogle · Gemini 3 Flash Preview#4Average Pass¹: 83.49%τ² standard agent
- τ² standard agent · Gemini 3 Pro Preview · highGoogle · Gemini 3 Pro Preview#5Average Pass¹: 82.46%τ² standard agent
- τ² standard agent · GLM-5 · thinkingZhipu AI · GLM-5#6Average Pass¹: 81.01%τ² standard agent
- τ² standard agent · Claude Sonnet 4.5 · extended thinkingAnthropic · Claude Sonnet 4.5#7Average Pass¹: 76.41%τ² standard agent
- τ² standard agent · GPT-5.2 · no reasoningOpenAI · GPT-5.2#8Average Pass¹: 61.58%τ² standard agent