OPENGPT EVALUATION METHODOLOGY

A model decision is only as credible as its evidence.

OpenGPT separates requirement matching, hands-on comparison, and published evidence. Every conclusion should state what was tested, under which conditions, and what remains unverified.

From a real task to a reviewable conclusion

1

Freeze the question

Record the task, success criteria, inputs, constraints, model version, date, and settings before comparing.

2

Run comparable trials

Use the same task and rubric for each candidate. Compare hides model labels in the review interface with a stable balanced order; this only reduces visible-label bias and is not double-blind or independent execution.

3

Keep the raw evidence

Retain prompts, outputs, scores, errors, latency, and hashes. Never replace missing evidence with an estimate.

4

Publish the limits

Report uncertainty, sample size, exclusions, conflicts, and what the result cannot prove.

Evidence-level target standard

Evidence v1/v2 checks format and declared consistency. Decision Workspace V3 additionally recomputes hashes, assignments, mappings, scores, and summaries from recorded inputs. It still does not establish L4 independent reproduction.

L1Supported now

Format checked

The package follows the declared schema and contains no obvious secret fields.

L2Partially supported

Protocol matched

Target: record the task suite, settings, model declaration, and evaluation date consistently. V1 currently checks only part of this declaration.

L3Supported for V3

Recomputable

Decision Workspace V3 includes raw pasted outputs and scoring inputs for strict local recomputation; Evidence v1/v2 does not carry the same fields.

L4Not yet supported

Independently reproduced

Target: a separate runner or signed source reproduces the result under comparable conditions. V1 does not verify this.

What local validation can establish

  • Required fields and supported data types
  • Internal consistency of scores, runs, dates, and hashes
  • Whether the package declares demo data
  • Absence of common API-key and secret fields

What it cannot establish alone

  • That the declared provider or model actually produced the output
  • That a score is free from selection bias or weak test design
  • That one result applies to every task, language, user, or future model version
  • That a model is safe, lawful, or appropriate for a high-stakes decision

Publication principles

No pay-to-rank

Sponsorship may be disclosed, but it cannot change scoring, evidence gates, or placement.

Versions, not brand names

A model family is not a stable test subject. Exact model IDs, settings, region, and date matter.

Real tasks before broad claims

OpenGPT recommends shortlists by requirements, then asks users to test their own representative work.

Uncertainty is a result

Ties, small samples, failed runs, and inconclusive evidence should remain visible.