OPENGPT EVIDENCE PROTOCOL V1

A model decision is only as credible as its evidence.

OpenGPT separates requirement matching, hands-on comparison, and published evidence. Every conclusion should state what was tested, under which conditions, and what remains unverified.

From a real task to a reviewable conclusion

1

Freeze the question

Record the task, success criteria, inputs, constraints, model version, date, and settings before comparing.

2

Run comparable trials

Use the same task and scoring rubric for each candidate. For blinded preference studies, use an external randomized runner; the current private comparison does not hide labels.

3

Keep the raw evidence

Retain prompts, outputs, scores, errors, latency, and hashes. Never replace missing evidence with an estimate.

4

Publish the limits

Report uncertainty, sample size, exclusions, conflicts, and what the result cannot prove.

Evidence-level target standard

The current local verifier fully supports L1, partially checks L2 declarations, and does not yet verify L3 or L4. Higher levels are protocol targets—not claims about uploaded files.

L1Supported now

Format checked

The package follows the declared schema and contains no obvious secret fields.

L2Partially supported

Protocol matched

Target: record the task suite, settings, model declaration, and evaluation date consistently. V1 currently checks only part of this declaration.

L3Not yet supported

Recomputable

Target: include raw outputs and scoring inputs so an independent party can recompute results. V1 does not yet carry these fields.

L4Not yet supported

Independently reproduced

Target: a separate runner or signed source reproduces the result under comparable conditions. V1 does not verify this.

What local validation can establish

  • Required fields and supported data types
  • Internal consistency of scores, runs, dates, and hashes
  • Whether the package declares demo data
  • Absence of common API-key and secret fields

What it cannot establish alone

  • That the declared provider or model actually produced the output
  • That a score is free from selection bias or weak test design
  • That one result applies to every task, language, user, or future model version
  • That a model is safe, lawful, or appropriate for a high-stakes decision

Publication principles

No pay-to-rank

Sponsorship may be disclosed, but it cannot change scoring, evidence gates, or placement.

Versions, not brand names

A model family is not a stable test subject. Exact model IDs, settings, region, and date matter.

Real tasks before broad claims

OpenGPT recommends shortlists by requirements, then asks users to test their own representative work.

Uncertainty is a result

Ties, small samples, failed runs, and inconclusive evidence should remain visible.