Agent configuration · τ³-Banking Knowledge

τ³ Banking AllTools · Gemini 2.5 Pro · high

The ranked object: model, scaffold, tools, permissions, budget, and environment together.

Pass¹13.66%
Cost per successful taskNot published
Published result rangeNot published

01

Configuration

Agent systemτ³ Banking AllTools · Gemini 2.5 Pro · high
Base modelGemini 2.5 Pro
Source model ID or labelGemini 2.5 Pro · high
Scaffoldτ³ standard agent
Scaffold versionv1.0.1
ToolsAllTools: BM25, text-embedding-3-large retrieval, and sandboxed shell
EnvironmentStandard scaffold with BM25, embedding retrieval, and sandboxed command-line search
Budget4 trials per task · 388 trajectories per configuration

02

Outcome

Pass¹13.66%
Partial scoreNot published
Cost per successful taskNot published
Reported benchmark costNot reported
Median completion timeNot published
Human interventionNot published

03

Operation

Repeat-run evidence4 runs
Published result rangeNot published
Safety noteNot published
Comparability grouptau3-banking-v1-0-1-alltools-gpt-5-2-low-seed-300-4-trials
Evaluation date
Data captured

04

Trust

EvidenceListed by the benchmark publisher
SourceSierra Research · τ³-Banking
Source snapshotv1.0.1 · AllTools · GPT-5.2 low user simulator · seed 300 · 4 trials · Jul 30, 2026
Last checkedJul 30, 2026
Tasks97 banking knowledge tasks
Evidence levelPublisher maintained
Evidence note

Included rows share benchmark v1.0.1, the 97-task Banking Knowledge split, AllTools setup, user simulator, seed, and four-trial protocol. Missing costs remain unknown.

Open audit

What this result does not prove

This result applies to the exact system and source conditions shown here. It does not establish a universal best agent, production reliability, data governance, or performance on a different benchmark version.

Compare like with like

Only configurations from the same benchmark version and task scope can be placed in one comparison.

Source-reported data; missing values remain unknown.

Validate before adoption

Public rankings create a shortlist. A small private trial decides whether the configuration fits your work.

  1. 1

    Choose 10–20 representative tasks and define a clear pass condition.

  2. 2

    Lock the agent, model, tools, permissions, and budget before testing.

  3. 3

    Run each important task more than once; record success, cost, time, and takeovers.

  4. 4

    Review failures and data-handling risks before a production decision.

All AI Agent rankings