Agent benchmarks

GAIA agent benchmark

OpenGPT keeps 15 published rows attached to the exact GAIA benchmark version and environment shown below.

Tasks
165
Published rows
15
Source data date
Identity reviewed
Historical

Evidence scope

HAL verified snapshot · Top 15 published configurations · public validation set · 165 questions. public validation questions. General-assistant questions requiring reasoning, multimodality, browsing, and tools. Source-specific; no normalized step or token limit published.

How OpenGPT handles this evidence

OpenGPT preserves the publisher's metric, task scope, environment, version, configuration fields, and source date. It does not create a cross-benchmark overall Agent score.

Comparison boundaries

  • A higher result applies only to the displayed benchmark version and conditions.
  • Agent results reflect the full system, not the base model alone.
  • Missing cost, reliability, safety, or latency data remains unreported rather than being inferred.
  • HAL reproduced these public-validation rows and publishes scaffold, model, score, cost, run count, and traces. The source has paused adding new models, and it does not publish a normalized tool or step budget.

HAL verified snapshot · Top 15 published configurations · public validation set · 165 questions

Showing 15 of 15 published rows. Every result remains attached to the source benchmark and configuration.

  1. HAL Generalist Agent · Claude Sonnet 4.5 (September 2025)Holistic Agent Leaderboard · Claude Sonnet 4.5 (September 2025)#1Accuracy: 74.55%HAL Generalist Agent
  2. HAL Generalist Agent · Claude Sonnet 4.5 High (September 2025)Holistic Agent Leaderboard · Claude Sonnet 4.5 High (September 2025)#2Accuracy: 70.91%HAL Generalist Agent
  3. HAL Generalist Agent · Claude Opus 4.1 High (August 2025)Holistic Agent Leaderboard · Claude Opus 4.1 High (August 2025)#3Accuracy: 68.48%HAL Generalist Agent
  4. HAL Generalist Agent · Claude Opus 4 High (May 2025)Holistic Agent Leaderboard · Claude Opus 4 High (May 2025)#4Accuracy: 64.85%HAL Generalist Agent
  5. HAL Generalist Agent · Claude-3.7 Sonnet High (February 2025)Holistic Agent Leaderboard · Claude-3.7 Sonnet High (February 2025)#5Accuracy: 64.24%HAL Generalist Agent
  6. HAL Generalist Agent · Claude Opus 4.1 (August 2025)Holistic Agent Leaderboard · Claude Opus 4.1 (August 2025)#6Accuracy: 64.24%HAL Generalist Agent
  7. HF Open Deep Research · GPT-5 Medium (August 2025)Hugging Face · GPT-5 Medium (August 2025)#7Accuracy: 62.80%HF Open Deep Research
  8. HAL Generalist Agent · GPT-5 Medium (August 2025)Holistic Agent Leaderboard · GPT-5 Medium (August 2025)#8Accuracy: 59.39%HAL Generalist Agent
  9. HAL Generalist Agent · o4-mini Low (April 2025)Holistic Agent Leaderboard · o4-mini Low (April 2025)#9Accuracy: 58.18%HAL Generalist Agent
  10. HF Open Deep Research · Claude Opus 4 (May 2025)Hugging Face · Claude Opus 4 (May 2025)#10Accuracy: 57.58%HF Open Deep Research
  11. HAL Generalist Agent · Claude-3.7 Sonnet (February 2025)Holistic Agent Leaderboard · Claude-3.7 Sonnet (February 2025)#11Accuracy: 56.36%HAL Generalist Agent
  12. HAL Generalist Agent · Claude Haiku 4.5 (October 2025)Holistic Agent Leaderboard · Claude Haiku 4.5 (October 2025)#12Accuracy: 56.36%HAL Generalist Agent
  13. HF Open Deep Research · o4-mini High (April 2025)Hugging Face · o4-mini High (April 2025)#13Accuracy: 55.76%HF Open Deep Research
  14. HAL Generalist Agent · o4-mini High (April 2025)Holistic Agent Leaderboard · o4-mini High (April 2025)#14Accuracy: 54.55%HAL Generalist Agent
  15. HF Open Deep Research · GPT-4.1 (April 2025)Hugging Face · GPT-4.1 (April 2025)#15Accuracy: 50.30%HF Open Deep Research