Agent configuration · GAIA
HAL Generalist Agent · Claude-3.7 Sonnet High (February 2025)
The ranked object: model, scaffold, tools, permissions, budget, and environment together.
01
Configuration
02
Outcome
03
Operation
04
Trust
Evidence note
HAL reproduced these public-validation rows and publishes the scaffold, model, score, cost, run count, and traces. Treat the snapshot as historical; no normalized tool or step budget is published.
Open audit ↗What this result does not prove
This result applies to the exact system and source conditions shown here. It does not establish a universal best agent, production reliability, data governance, or performance on a different benchmark version.
Source-reported data; missing values remain unknown.
Validate before adoption
Public rankings create a shortlist. A small private trial decides whether the configuration fits your work.
- 1
Choose 10–20 representative tasks and define a clear pass condition.
- 2
Lock the agent, model, tools, permissions, and budget before testing.
- 3
Run each important task more than once; record success, cost, time, and takeovers.
- 4
Review failures and data-handling risks before a production decision.