Agent configuration · Berkeley Function-Calling Leaderboard

BFCL V4 · Grok-4-1-fast-reasoning (FC)

The ranked object: model, scaffold, tools, permissions, budget, and environment together.

Overall accuracy69.57%
Cost per successful taskNot published
Published result rangeNot published

01

Configuration

Agent systemBFCL V4 · Grok-4-1-fast-reasoning (FC)
Base modelGrok-4-1-fast-reasoning (FC)
Source model ID or labelGrok-4-1-fast-reasoning (FC)
ScaffoldBFCL V4 tool-use evaluation
Scaffold versionbfcl-eval 2025.12.17
ToolsNative function calling
EnvironmentFunction calling, web search, memory, multi-turn, hallucination, and format-sensitivity tests
BudgetCategory-specific publisher protocol

02

Outcome

Overall accuracy69.57%
Partial scoreNot published
Cost per successful taskNot published
Reported benchmark cost$17.26
Median completion timeNot published
Human interventionNot published

03

Operation

Repeat-run evidenceNot published
Published result rangeNot published
Safety noteNot published
Comparability groupbfcl-v4-f7cf735-2025-12-17-overall
Evaluation date
Data captured

04

Trust

EvidenceListed by the benchmark publisher
SourceBFCL V4
Source snapshotBFCL V4 · commit f7cf735 · bfcl-eval 2025.12.17 · Apr 12, 2026
Last checkedAug 1, 2026
TasksNot published
Evidence levelPublisher maintained
Evidence note

Includes all 109 publisher rows from one frozen V4 CSV. All rows share commit f7cf735, bfcl-eval 2025.12.17, and the same overall-accuracy composition.

Open audit

What this result does not prove

This result applies to the exact system and source conditions shown here. It does not establish a universal best agent, production reliability, data governance, or performance on a different benchmark version.

Compare like with like

Only configurations from the same benchmark version and task scope can be placed in one comparison.

Source-reported data; missing values remain unknown.

Validate before adoption

Public rankings create a shortlist. A small private trial decides whether the configuration fits your work.

  1. 1

    Choose 10–20 representative tasks and define a clear pass condition.

  2. 2

    Lock the agent, model, tools, permissions, and budget before testing.

  3. 3

    Run each important task more than once; record success, cost, time, and takeovers.

  4. 4

    Review failures and data-handling risks before a production decision.

All AI Agent rankings