Agent benchmarks

WebArena agent benchmark

OpenGPT keeps 0 published rows attached to the exact WebArena benchmark version and environment shown below.

Tasks
812
Published rows
0
Source data date
Identity reviewed
Review needed

Evidence scope

v0.2.0 · canonical self-hosted environment. realistic web tasks. Self-hosted websites with functional outcome evaluators. Submission-specific.

How OpenGPT handles this evidence

OpenGPT preserves the publisher's metric, task scope, environment, version, configuration fields, and source date. It does not create a cross-benchmark overall Agent score.

Comparison boundaries

  • A higher result applies only to the displayed benchmark version and conditions.
  • Agent results reflect the full system, not the base model alone.
  • Missing cost, reliability, safety, or latency data remains unreported rather than being inferred.
  • OpenGPT has verified the official benchmark scope, but has not yet frozen a configuration-complete, directly comparable result snapshot.

v0.2.0 · canonical self-hosted environment

Showing 0 of 0 published rows. Every result remains attached to the source benchmark and configuration.