TRUST & DATA STATUS

What is checked, reviewed, and published

See the evidence status behind OpenGPT's model and Agent rankings, including source dates, latest checks, and publication controls.

Status policy updated

01

Official model lifecycle verification

Lifecycle labels are recorded only when an official provider source supports the exact model identity. A daily read-only check flags missing evidence and aging reviews.

Stale after more than 30 days79 exact model identities tracked78 lifecycle records are stale and must be rechecked against the official source.
Latest lifecycle reviewDeepSeek V4 Prodeepseek-v4-pro
Last verified
Review due
Official lifecycle source
Open source
General availability
78 lifecycle records are stale and must be rechecked against the official source.

02

Model ranking sources

Live monitoring states come from the public status endpoint. Source dates remain separate from check times.

SourceEvidence stateSource data dateCheck activityPublication control
Reading live state
Latest attemptLast successfulNot published by source
Automatic after validation
Reading live state
Latest attemptLast successfulNot published by source
Automatic after validation

03

Agent evaluation sources

Agent rows show the published evidence record and its latest recorded check. They do not claim source or product uptime.

SourceEvidence stateSource data dateCheck activityPublication control
SWE-bench VerifiedOpen source
Historical snapshot
Latest attemptLast successfulNot published by source
Excluded until adapter validation
Terminal-Bench 2.1Open source
Published snapshot
Latest attemptLast successfulNot published by source
Excluded until adapter validation
AssistantBenchOpen source
Historical snapshot
Not published by source
Latest attemptLast successfulNot published by source
Excluded until adapter validation
Historical snapshot
Not published by source
Latest attemptLast successfulNot published by source
Excluded until adapter validation
OSWorld 2.0Open source
Historical snapshot
Latest attemptLast successfulNot published by source
Excluded until adapter validation
Berkeley Function-Calling LeaderboardOpen source
Historical snapshot
Latest attemptLast successfulNot published by source
Excluded until adapter validation
Historical snapshot
Not published by source
Latest attemptLast successfulNot published by source
Excluded until adapter validation
τ²-bench CoreOpen source
Published snapshot
Latest attemptLast successfulNot published by source
Excluded until adapter validation
τ³-Banking KnowledgeOpen source
Published snapshot
Latest attemptLast successfulNot published by source
Excluded until adapter validation

04

Service boundaries

A status label only describes the specific workflow named here. It should not be read as a broader availability promise.

Public rankings

OpenGPT publishes source snapshots and catalog records. It does not continuously test provider APIs, model endpoints, prices, or regional access.

Source monitoring

Configured upstream artifacts are checked on their stated cadence. A successful check does not prove every source page, model, or API is available.

Local decision tools

Saved candidates, task inputs, and reports stay in this browser by default. This status page does not imply cloud sync or recovery.

Submissions and email

Submissions enter editorial review. Receipt and final-result emails use a queued transactional process; instant delivery or a response time is not promised.

05

How publication decisions work

Collection, review, catalog inclusion, comparability, and ranking remain separate decisions.

01

Validated automatic publication

LMArena and LiveBench are eligible only after their versioned schema, completeness, date, integrity, and anomaly checks pass. Validation fails closed.

02

Adapter coverage

Structured Agent sources publish only through validated adapters. Sources without a safe adapter remain paused or historical and are not used as current evidence.

03

User submissions

Approval can list a reviewed Agent product in the public directory. It does not set rank, comparability, evidence grade, or guarantee leaderboard inclusion.

06

Methods and corrections

Read the comparison rules or report an inaccurate date, mapping, link, or evidence label.