A model decision is only as credible as its evidence.
OpenGPT separates requirement matching, hands-on comparison, and published evidence. Every conclusion should state what was tested, under which conditions, and what remains unverified.
From a real task to a reviewable conclusion
Freeze the question
Record the task, success criteria, inputs, constraints, model version, date, and settings before comparing.
Run comparable trials
Use the same task and rubric for each candidate. Compare hides model labels in the review interface with a stable balanced order; this only reduces visible-label bias and is not double-blind or independent execution.
Keep the raw evidence
Retain prompts, outputs, scores, errors, latency, and hashes. Never replace missing evidence with an estimate.
Publish the limits
Report uncertainty, sample size, exclusions, conflicts, and what the result cannot prove.
Evidence-level target standard
Evidence v1/v2 checks format and declared consistency. Decision Workspace V3 additionally recomputes hashes, assignments, mappings, scores, and summaries from recorded inputs. It still does not establish L4 independent reproduction.
Format checked
The package follows the declared schema and contains no obvious secret fields.
Protocol matched
Target: record the task suite, settings, model declaration, and evaluation date consistently. V1 currently checks only part of this declaration.
Recomputable
Decision Workspace V3 includes raw pasted outputs and scoring inputs for strict local recomputation; Evidence v1/v2 does not carry the same fields.
Independently reproduced
Target: a separate runner or signed source reproduces the result under comparable conditions. V1 does not verify this.
What local validation can establish
- ✓Required fields and supported data types
- ✓Internal consistency of scores, runs, dates, and hashes
- ✓Whether the package declares demo data
- ✓Absence of common API-key and secret fields
What it cannot establish alone
- —That the declared provider or model actually produced the output
- —That a score is free from selection bias or weak test design
- —That one result applies to every task, language, user, or future model version
- —That a model is safe, lawful, or appropriate for a high-stakes decision
Publication principles
No pay-to-rank
Sponsorship may be disclosed, but it cannot change scoring, evidence gates, or placement.
Versions, not brand names
A model family is not a stable test subject. Exact model IDs, settings, region, and date matter.
Real tasks before broad claims
OpenGPT recommends shortlists by requirements, then asks users to test their own representative work.
Uncertainty is a result
Ties, small samples, failed runs, and inconclusive evidence should remain visible.