Agent評測

τ²-bench Core Agent評測

OpenGPT將8筆公開記錄保留在下方所示的 τ²-bench Core評測版本與環境內。

工作
278
公開記錄
8
來源資料日期
身分核驗日期
目前

資料範圍

v1.0.1 · standard agent · GPT-5.2 low user simulator · 4 trials. base tasks across retail, airline, and telecom. Standard agent, standard domain tools, and a shared GPT-5.2 low-reasoning user simulator. 4 trials per task · 1,112 trajectories per configuration.

OpenGPT如何處理這些資料

OpenGPT保留發布方的指標、工作範圍、環境、版本、設定欄位與資料日期,不產生跨評測Agent總分。

可比範圍

  • 更高結果只適用於所展示的評測版本與條件。
  • Agent結果反映完整系統,而不是基礎模型單獨能力。
  • 成本、穩定性、安全性或速度沒有公開時保持未報告,不進行推測。
  • OpenGPT 僅收錄了八條由發佈方提交、且共同使用基準 v1.0.1、標準框架和工具、相同使用者模擬器、相同基礎任務劃分以及四次試驗的記錄。其他 τ-bench 記錄不在此範圍內。

v1.0.1 · standard agent · GPT-5.2 low user simulator · 4 trials

展示8筆公開記錄中的8筆;每項結果都保留來源評測與設定。

  1. τ² standard agent · Qwen3.5-397B-A17B · thinkingAlibaba Cloud · Qwen3.5-397B-A17B#1Average Pass¹: 87.91%τ² standard agent
  2. τ² standard agent · Claude Opus 4.5 · 高Anthropic · Claude Opus 4.5#2Average Pass¹: 85.31%τ² standard agent
  3. τ² standard agent · GPT-5.2 · 高OpenAI · GPT-5.2#3Average Pass¹: 84.76%τ² standard agent
  4. τ² standard agent · Gemini 3 Flash Preview · 高Google · Gemini 3 Flash Preview#4Average Pass¹: 83.49%τ² standard agent
  5. τ² standard agent · Gemini 3 Pro Preview · 高Google · Gemini 3 Pro Preview#5Average Pass¹: 82.46%τ² standard agent
  6. τ² standard agent · GLM-5 · thinkingZhipu AI · GLM-5#6Average Pass¹: 81.01%τ² standard agent
  7. τ² standard agent · Claude Sonnet 4.5 · 擴展思考Anthropic · Claude Sonnet 4.5#7Average Pass¹: 76.41%τ² standard agent
  8. τ² standard agent · GPT-5.2 · 無推理OpenAI · GPT-5.2#8Average Pass¹: 61.58%τ² standard agent