Agent configuration · AssistantBench
Browser-Use · Claude-3.7 Sonnet (February 2025)
The ranked object: model, scaffold, tools, permissions, budget, and environment together.
01
Configuration
02
Outcome
03
Operation
04
Trust
Evidence note
HAL reproduced these public rows and publishes traces, costs, and run counts. Treat the snapshot as historical: HAL paused new models, and reported costs exclude caching benefits.
Open audit ↗What this result does not prove
This result applies to the exact system and source conditions shown here. It does not establish a universal best agent, production reliability, data governance, or performance on a different benchmark version.
Source-reported data; missing values remain unknown.
Validate before adoption
Public rankings create a shortlist. A small private trial decides whether the configuration fits your work.
- 1
Choose 10–20 representative tasks and define a clear pass condition.
- 2
Lock the agent, model, tools, permissions, and budget before testing.
- 3
Run each important task more than once; record success, cost, time, and takeovers.
- 4
Review failures and data-handling risks before a production decision.