Agent configuration · τ³-Banking Knowledge

τ³ Banking AllTools · Claude Opus 4.7 · max

The ranked object: model, scaffold, tools, permissions, budget, and environment together.

Pass¹30.15%
Cost per successful taskNot published
Published result rangeNot published

01

Configuration

Agent systemτ³ Banking AllTools · Claude Opus 4.7 · max
Base modelClaude Opus 4.7
Source model ID or labelClaude Opus 4.7 · max
Scaffoldτ³ standard agent
Scaffold versionv1.0.1
ToolsAllTools: BM25, text-embedding-3-large retrieval, and sandboxed shell
EnvironmentStandard scaffold with BM25, embedding retrieval, and sandboxed command-line search
Budget4 trials per task · 388 trajectories per configuration

02

Outcome

Pass¹30.15%
Partial scoreNot published
Cost per successful taskNot published
Reported benchmark costNot reported
Median completion timeNot published
Human interventionNot published

03

Operation

Repeat-run evidence4 runs
Published result rangeNot published
Safety noteNot published
Comparability grouptau3-banking-v1-0-1-alltools-gpt-5-2-low-seed-300-4-trials
Evaluation date
Data captured

04

Trust

EvidenceListed by the benchmark publisher
SourceSierra Research · τ³-Banking
Source snapshotv1.0.1 · AllTools · GPT-5.2 low user simulator · seed 300 · 4 trials
Source data dateNot published by source
Retrieved by OpenGPT
Ranking last changedNo comparable history yet
Last checked
Tasks97 banking knowledge tasks
Evidence levelPublisher maintained
Evidence note

Included rows share benchmark v1.0.1, the 97-task Banking Knowledge split, AllTools setup, user simulator, seed, and four-trial protocol. Missing costs remain unknown.

Open audit

What this result does not prove

This result applies to the exact system and source conditions shown here. It does not establish a universal best agent, production reliability, data governance, or performance on a different benchmark version.

Compare like with like

Only configurations from the same benchmark version and task scope can be placed in one comparison.

Source-reported data; missing values remain unknown.

Validate before adoption

Public rankings create a shortlist. A small private trial decides whether the configuration fits your work.

  1. 1

    Choose 10–20 representative tasks and define a clear pass condition.

  2. 2

    Lock the agent, model, tools, permissions, and budget before testing.

  3. 3

    Run each important task more than once; record success, cost, time, and takeovers.

  4. 4

    Review failures and data-handling risks before a production decision.

All AI Agent rankings