Tool-calling model signal · Berkeley Function-Calling Leaderboard

BFCL V4 · Gemini-3-Pro-Preview (Prompt)

This BFCL result describes a model and function-calling setup, not a full Agent product or production-readiness result.

Overall accuracy72.51%
Cost per successful taskNot published
Published result rangeNot published

01

Configuration

Agent systemBFCL V4 · Gemini-3-Pro-Preview (Prompt)
Base modelGemini-3-Pro-Preview (Prompt)
Source model ID or labelGemini-3-Pro-Preview (Prompt)
ScaffoldBFCL V4 tool-use evaluation
Scaffold versionbfcl-eval 2025.12.17
ToolsPrompt-based function calling
EnvironmentFunction calling, web search, memory, multi-turn, hallucination, and format-sensitivity tests
BudgetCategory-specific publisher protocol

02

Outcome

Overall accuracy72.51%
Partial scoreNot published
Cost per successful taskNot published
Reported benchmark cost$298.47
Median completion timeNot published
Human interventionNot published

03

Operation

Repeat-run evidenceNot published
Published result rangeNot published
Safety noteNot published
Comparability groupbfcl-v4-f7cf735-2025-12-17-overall
Evaluation date
Data captured

04

Trust

EvidenceListed by the benchmark publisher
SourceBFCL V4
Source snapshotBFCL V4 · commit f7cf735 · bfcl-eval 2025.12.17
Source data dateNot published by source
Retrieved by OpenGPT
Ranking last changedNo comparable history yet
Last checked
TasksNot published
Evidence levelPublisher maintained
Evidence note

Includes all 109 publisher rows from one frozen V4 CSV. All rows share commit f7cf735, bfcl-eval 2025.12.17, and the same overall-accuracy composition.

Open audit

What this result does not prove

This BFCL result describes a model and function-calling setup, not a full Agent product or production-readiness result.

Compare like with like

Only configurations from the same benchmark version and task scope can be placed in one comparison.

Source-reported data; missing values remain unknown.

Validate before adoption

Public rankings create a shortlist. A small private trial decides whether the configuration fits your work.

  1. 1

    Choose 10–20 representative tasks and define a clear pass condition.

  2. 2

    Lock the agent, model, tools, permissions, and budget before testing.

  3. 3

    Run each important task more than once; record success, cost, time, and takeovers.

  4. 4

    Review failures and data-handling risks before a production decision.

All AI Agent rankings