Shared benchmark scope · τ³-Banking Knowledge

Same-benchmark agent comparison

The entries below share the same benchmark scope. Their exact scaffolds, model versions, dates, and original source values remain visible.

A higher source value only means better performance on this benchmark and configuration. It is not a cross-benchmark or production-readiness score.

Compare like with like

Decision overview

Only configurations from the same benchmark version and task scope can be placed in one comparison.

Source-reported data; missing values remain unknown.
Thinking Machines Labτ³ Banking AllTools · Inkling · max
Pass¹25.00%Cost per successful taskNot publishedPublished result rangeNot published
Agent details
Qwenτ³ Banking AllTools · Qwen 3.8 Max · xhigh
Pass¹55.15%Cost per successful taskNot publishedPublished result rangeNot published
Agent details
17 identical or entirely unreported fields hidden
Agent systemτ³ Banking AllTools · Inkling · maxτ³ Banking AllTools · Qwen 3.8 Max · xhigh
Configuration
Base modelInklingQwen 3.8 Max
Source model ID or labelInkling · maxQwen 3.8 Max · xhigh
Outcome
Pass¹25.00%55.15%
Trust
Evaluation dateJul 24, 2026Aug 3, 2026
Source
Open sourceOpen sourceOpen source

Configuration

τ³ Banking AllTools · Inkling · max

Base model
Inkling
Source model ID or label
Inkling · max

τ³ Banking AllTools · Qwen 3.8 Max · xhigh

Base model
Qwen 3.8 Max
Source model ID or label
Qwen 3.8 Max · xhigh

Outcome

τ³ Banking AllTools · Inkling · max

Pass¹
25.00%

τ³ Banking AllTools · Qwen 3.8 Max · xhigh

Pass¹
55.15%

Trust

τ³ Banking AllTools · Inkling · max

Evaluation date
Jul 24, 2026

τ³ Banking AllTools · Qwen 3.8 Max · xhigh

Evaluation date
Aug 3, 2026

Source

Source-reported data; missing values remain unknown.

Validate before adoption

Public rankings create a shortlist. A small private trial decides whether the configuration fits your work.

  1. 1

    Choose 10–20 representative tasks and define a clear pass condition.

  2. 2

    Lock the agent, model, tools, permissions, and budget before testing.

  3. 3

    Run each important task more than once; record success, cost, time, and takeovers.

  4. 4

    Review failures and data-handling risks before a production decision.

All AI Agent rankings