Agent configuration · Terminal-Bench 2.1
Codex · GPT-5.6 Terra · max
The ranked object: model, scaffold, tools, permissions, budget, and environment together.
01
Configuration
02
Outcome
03
Operation
04
Trust
Evidence note
The Terminal-Bench team says it ran and verified these submissions. Its 2.1 release revised tasks and the evaluation environment.
Open audit ↗What this result does not prove
This result applies to the exact system and source conditions shown here. It does not establish a universal best agent, production reliability, data governance, or performance on a different benchmark version.
Source-reported data; missing values remain unknown.
Validate before adoption
Public rankings create a shortlist. A small private trial decides whether the configuration fits your work.
- 1
Choose 10–20 representative tasks and define a clear pass condition.
- 2
Lock the agent, model, tools, permissions, and budget before testing.
- 3
Run each important task more than once; record success, cost, time, and takeovers.
- 4
Review failures and data-handling risks before a production decision.