Agent benchmarks
WebArena agent benchmark
OpenGPT keeps 0 published rows attached to the exact WebArena benchmark version and environment shown below.
- Tasks
- 812
- Published rows
- 0
- Source data date
- Identity reviewed
- Review needed
Evidence scope
v0.2.0 · canonical self-hosted environment. realistic web tasks. Self-hosted websites with functional outcome evaluators. Submission-specific.
How OpenGPT handles this evidence
OpenGPT preserves the publisher's metric, task scope, environment, version, configuration fields, and source date. It does not create a cross-benchmark overall Agent score.
Comparison boundaries
- A higher result applies only to the displayed benchmark version and conditions.
- Agent results reflect the full system, not the base model alone.
- Missing cost, reliability, safety, or latency data remains unreported rather than being inferred.
- OpenGPT has verified the official benchmark scope, but has not yet frozen a configuration-complete, directly comparable result snapshot.
v0.2.0 · canonical self-hosted environment
Showing 0 of 0 published rows. Every result remains attached to the source benchmark and configuration.