A model decision is only as credible as its evidence.
OpenGPT separates requirement matching, hands-on comparison, and published evidence. Every conclusion should state what was tested, under which conditions, and what remains unverified.
From a real task to a reviewable conclusion
Freeze the question
Record the task, success criteria, inputs, constraints, model version, date, and settings before comparing.
Run comparable trials
Use the same task and scoring rubric for each candidate. For blinded preference studies, use an external randomized runner; the current private comparison does not hide labels.
Keep the raw evidence
Retain prompts, outputs, scores, errors, latency, and hashes. Never replace missing evidence with an estimate.
Publish the limits
Report uncertainty, sample size, exclusions, conflicts, and what the result cannot prove.
Evidence-level target standard
The current local verifier fully supports L1, partially checks L2 declarations, and does not yet verify L3 or L4. Higher levels are protocol targets—not claims about uploaded files.
Format checked
The package follows the declared schema and contains no obvious secret fields.
Protocol matched
Target: record the task suite, settings, model declaration, and evaluation date consistently. V1 currently checks only part of this declaration.
Recomputable
Target: include raw outputs and scoring inputs so an independent party can recompute results. V1 does not yet carry these fields.
Independently reproduced
Target: a separate runner or signed source reproduces the result under comparable conditions. V1 does not verify this.
What local validation can establish
- ✓Required fields and supported data types
- ✓Internal consistency of scores, runs, dates, and hashes
- ✓Whether the package declares demo data
- ✓Absence of common API-key and secret fields
What it cannot establish alone
- —That the declared provider or model actually produced the output
- —That a score is free from selection bias or weak test design
- —That one result applies to every task, language, user, or future model version
- —That a model is safe, lawful, or appropriate for a high-stakes decision
Publication principles
No pay-to-rank
Sponsorship may be disclosed, but it cannot change scoring, evidence gates, or placement.
Versions, not brand names
A model family is not a stable test subject. Exact model IDs, settings, region, and date matter.
Real tasks before broad claims
OpenGPT recommends shortlists by requirements, then asks users to test their own representative work.
Uncertainty is a result
Ties, small samples, failed runs, and inconclusive evidence should remain visible.