TRUST & DATA STATUS

What is checked, reviewed, and published

See the evidence status behind OpenGPT's model and Agent rankings, including source dates, latest checks, and publication controls.

Status policy updated

01

Model ranking sources

Live monitoring states come from the public status endpoint. Source dates remain separate from check times.

SourceEvidence stateSource data dateLast checkedPublication control
Reading live state
Automatic after validation
Artificial AnalysisOpen source
Reading live state
Not published by source
Human review before publication
Reading live state
Human review before publication

02

Agent evaluation sources

Agent rows show the published evidence record and its latest recorded check. They do not claim source or product uptime.

SourceEvidence stateSource data dateLast checkedPublication control
SWE-bench VerifiedOpen source
Historical snapshot
Human review before publication
Terminal-Bench 2.1Open source
Published snapshot
Human review before publication
AssistantBenchOpen source
Historical snapshot
Not published by source
Human review before publication
Source review required
Not published by source
Human review before publication
OSWorld 2.0Open source
Published snapshot
Human review before publication
Berkeley Function-Calling LeaderboardOpen source
Historical snapshot
Human review before publication
Historical snapshot
Not published by source
Human review before publication
τ²-bench CoreOpen source
Published snapshot
Not published by source
Human review before publication
τ³-Banking KnowledgeOpen source
Published snapshot
Human review before publication

03

Service boundaries

A status label only describes the specific workflow named here. It should not be read as a broader availability promise.

Public rankings

OpenGPT publishes source snapshots and catalog records. It does not continuously test provider APIs, model endpoints, prices, or regional access.

Source monitoring

Configured upstream artifacts are checked on their stated cadence. A successful check does not prove every source page, model, or API is available.

Local decision tools

Saved candidates, task inputs, and reports stay in this browser by default. This status page does not imply cloud sync or recovery.

Submissions and email

Submissions enter editorial review. Receipt and final-result emails use a queued transactional process; instant delivery or a response time is not promised.

04

How publication decisions work

Collection, review, catalog inclusion, comparability, and ranking remain separate decisions.

01

Validated automatic publication

Only LMArena is eligible, and only after schema, completeness, date, and anomaly checks pass. Validation fails closed.

02

Human source review

Artificial Analysis, LiveBench, and Agent sources require human review before published evidence changes.

03

User submissions

Approval can list a reviewed Agent product in the public directory. It does not set rank, comparability, evidence grade, or guarantee leaderboard inclusion.

05

Methods and corrections

Read the comparison rules or report an inaccurate date, mapping, link, or evidence label.