Agent评测
τ²-bench Core Agent评测
OpenGPT将8条公开记录保留在下方所示的 τ²-bench Core评测版本与环境内。
- 任务
- 278
- 公开记录
- 8
- 来源数据日期
- 身份核验日期
- 当前
数据范围
v1.0.1 · standard agent · GPT-5.2 low user simulator · 4 trials. base tasks across retail, airline, and telecom. Standard agent, standard domain tools, and a shared GPT-5.2 low-reasoning user simulator. 4 trials per task · 1,112 trajectories per configuration.
OpenGPT如何处理这些数据
OpenGPT保留发布方的指标、任务范围、环境、版本、配置字段与数据日期,不生成跨评测Agent总分。
可比边界
- 更高结果只适用于所展示的评测版本与条件。
- Agent结果反映完整系统,而不是基础模型单独能力。
- 成本、稳定性、安全性或速度没有公开时保持未报告,不进行推测。
- OpenGPT 仅收录了八条由发布方提交、且共同使用基准 v1.0.1、标准框架和工具、相同用户模拟器、相同基础任务划分以及四次试验的记录。其他 τ-bench 记录不在此范围内。
v1.0.1 · standard agent · GPT-5.2 low user simulator · 4 trials
展示8条公开记录中的8条;每项结果都保留来源评测与配置。
- τ² standard agent · Qwen3.5-397B-A17B · thinkingAlibaba Cloud · Qwen3.5-397B-A17B#1Average Pass¹: 87.91%τ² standard agent
- τ² standard agent · Claude Opus 4.5 · 高Anthropic · Claude Opus 4.5#2Average Pass¹: 85.31%τ² standard agent
- τ² standard agent · GPT-5.2 · 高OpenAI · GPT-5.2#3Average Pass¹: 84.76%τ² standard agent
- τ² standard agent · Gemini 3 Flash Preview · 高Google · Gemini 3 Flash Preview#4Average Pass¹: 83.49%τ² standard agent
- τ² standard agent · Gemini 3 Pro Preview · 高Google · Gemini 3 Pro Preview#5Average Pass¹: 82.46%τ² standard agent
- τ² standard agent · GLM-5 · thinkingZhipu AI · GLM-5#6Average Pass¹: 81.01%τ² standard agent
- τ² standard agent · Claude Sonnet 4.5 · 扩展思考Anthropic · Claude Sonnet 4.5#7Average Pass¹: 76.41%τ² standard agent
- τ² standard agent · GPT-5.2 · 无推理OpenAI · GPT-5.2#8Average Pass¹: 61.58%τ² standard agent