Agent configuration · AssistantBench

Browser-Use · Gemini 2.0 Flash (February 2025)

The ranked object: model, scaffold, tools, permissions, budget, and environment together.

Accuracy2.62%
Cost per successful task$2.52
Published result rangeNot published

01

Configuration

Agent systemBrowser-Use · Gemini 2.0 Flash (February 2025)
Base modelGemini 2.0 Flash (February 2025)
Source model ID or labelGemini 2.0 Flash (February 2025)
ScaffoldBrowser-Use
Scaffold versionNot reported
ToolsBrowser-Use live-web browser automation
EnvironmentLive-web research and browsing tasks reproduced by the Holistic Agent Leaderboard
BudgetSource-specific · no normalized step or token limit published

02

Outcome

Accuracy2.62%
Partial scoreNot published
Cost per successful task$2.52
Reported benchmark cost$2.18
Median completion timeNot published
Human interventionNot published

03

Operation

Repeat-run evidence1 runs
Published result rangeNot published
Safety noteNot published
Comparability grouphal-assistantbench-public-33-browser-use
Evaluation dateNot published
Data captured

04

Trust

EvidenceComplete public run record
SourceHolistic Agent Leaderboard · AssistantBench
Source snapshotHAL verified snapshot · 33 public tasks · 1 scaffold · Jul 30, 2026
Last checkedJul 30, 2026
Tasks33 public, automatically verifiable browsing tasks
Evidence levelExternal caution
Evidence note

HAL reproduced these public rows and publishes traces, costs, and run counts. Treat the snapshot as historical: HAL paused new models, and reported costs exclude caching benefits.

Open audit

What this result does not prove

This result applies to the exact system and source conditions shown here. It does not establish a universal best agent, production reliability, data governance, or performance on a different benchmark version.

Compare like with like

Only configurations from the same benchmark version and task scope can be placed in one comparison.

Source-reported data; missing values remain unknown.

Validate before adoption

Public rankings create a shortlist. A small private trial decides whether the configuration fits your work.

  1. 1

    Choose 10–20 representative tasks and define a clear pass condition.

  2. 2

    Lock the agent, model, tools, permissions, and budget before testing.

  3. 3

    Run each important task more than once; record success, cost, time, and takeovers.

  4. 4

    Review failures and data-handling risks before a production decision.

All AI Agent rankings