Agent configuration · Terminal-Bench 2.1

Claude Code · Fable 5 · xhigh

The ranked object: model, scaffold, tools, permissions, budget, and environment together.

Task accuracy83.8%
Cost per successful taskNot published
Published result range±1.2%

01

Configuration

Agent systemClaude Code · Fable 5 · xhigh
Base modelFable 5
Source model ID or labelFable 5
ScaffoldClaude Code
Scaffold versionNot reported
ToolsTerminal shell
EnvironmentHarbor-managed terminal environments
BudgetFixed limits · at least 5 trials per task

02

Outcome

Task accuracy83.8%
Partial scoreNot published
Cost per successful taskNot published
Reported benchmark cost$552.67
Median completion timeNot published
Human interventionNot published

03

Operation

Repeat-run evidenceNot published
Published result range±1.2%
Safety noteNot published
Comparability groupterminal-bench-2-1-main
Evaluation date
Data captured

04

Trust

EvidenceListed by the benchmark publisher
SourceTerminal-Bench 2.1
Source snapshotTerminal-Bench 2.1 · publisher-verified leaderboard · Jul 30, 2026
Last checkedJul 30, 2026
Tasks89 terminal tasks
Evidence levelPublisher maintained
Evidence note

The Terminal-Bench team says it ran and verified these submissions. Its 2.1 release revised tasks and the evaluation environment.

Open audit

What this result does not prove

This result applies to the exact system and source conditions shown here. It does not establish a universal best agent, production reliability, data governance, or performance on a different benchmark version.

Compare like with like

Only configurations from the same benchmark version and task scope can be placed in one comparison.

Source-reported data; missing values remain unknown.

Validate before adoption

Public rankings create a shortlist. A small private trial decides whether the configuration fits your work.

  1. 1

    Choose 10–20 representative tasks and define a clear pass condition.

  2. 2

    Lock the agent, model, tools, permissions, and budget before testing.

  3. 3

    Run each important task more than once; record success, cost, time, and takeovers.

  4. 4

    Review failures and data-handling risks before a production decision.

All AI Agent rankings