에이전트 벤치마크

τ²-bench Core 에이전트 벤치마크

τ²-bench Core의 표시된 버전과 환경에 연결된 공개 행 8개를 유지합니다.

과제
278
공개 행
8
출처 데이터 날짜
ID 검토일
현재

근거 범위

v1.0.1 · standard agent · GPT-5.2 low user simulator · 4 trials. base tasks across retail, airline, and telecom. Standard agent, standard domain tools, and a shared GPT-5.2 low-reasoning user simulator. 4 trials per task · 1,112 trajectories per configuration.

OpenGPT의 근거 처리 방식

게시자의 지표, 과제 범위, 환경, 버전, 구성과 날짜를 보존하며 벤치마크 간 종합 점수를 만들지 않습니다.

비교 범위

  • 높은 결과는 표시된 벤치마크 버전과 조건에만 적용됩니다.
  • 에이전트 결과는 기반 모델 하나가 아닌 전체 시스템을 반영합니다.
  • 비용, 신뢰성, 안전성 또는 지연 시간 누락값은 추정하지 않습니다.
  • OpenGPT includes only the eight publisher-submitted rows that share benchmark v1.0.1, the standard scaffold and tools, the same user simulator, the same base task splits, and four trials. Other τ-bench rows remain outside this scope.

v1.0.1 · standard agent · GPT-5.2 low user simulator · 4 trials

공개 8개 중 8개를 표시합니다. 모든 결과는 출처 벤치마크와 구성에 연결됩니다.

  1. τ² standard agent · Qwen3.5-397B-A17B · thinkingAlibaba Cloud · Qwen3.5-397B-A17B#1Average Pass¹: 87.91%τ² standard agent
  2. τ² standard agent · Claude Opus 4.5 · highAnthropic · Claude Opus 4.5#2Average Pass¹: 85.31%τ² standard agent
  3. τ² standard agent · GPT-5.2 · highOpenAI · GPT-5.2#3Average Pass¹: 84.76%τ² standard agent
  4. τ² standard agent · Gemini 3 Flash Preview · highGoogle · Gemini 3 Flash Preview#4Average Pass¹: 83.49%τ² standard agent
  5. τ² standard agent · Gemini 3 Pro Preview · highGoogle · Gemini 3 Pro Preview#5Average Pass¹: 82.46%τ² standard agent
  6. τ² standard agent · GLM-5 · thinkingZhipu AI · GLM-5#6Average Pass¹: 81.01%τ² standard agent
  7. τ² standard agent · Claude Sonnet 4.5 · extended thinkingAnthropic · Claude Sonnet 4.5#7Average Pass¹: 76.41%τ² standard agent
  8. τ² standard agent · GPT-5.2 · no reasoningOpenAI · GPT-5.2#8Average Pass¹: 61.58%τ² standard agent