GUIDES & SUPPORT

OpenGPT Help Center

Choose a guide, follow the steps, or search for a problem.

Start here: three different jobs

Choosing, comparing, and verifying support different decisions. None of them is a universal leaderboard.

1ChooseNarrow model families from your requirements.2CompareCompare real outputs from exact versions on the same tasks.3VerifyCheck the structure and internal consistency of an evidence file.

01

Choose an AI model

Turn four practical requirements into an explainable shortlist.

Open this feature

What this feature is for

Use the finder when you know the work you need to do but do not know which model family to investigate first.

Before you start

Know the main task, where the model may run, your most important operating priority, and any language requirement.

Step-by-step tutorial

  1. 1Choose the task closest to your real work.
  2. 2Set the deployment boundary, especially when data cannot leave your environment.
  3. 3Choose the priority you are least willing to compromise.
  4. 4Set the language need, then review the match reasons, risks, and validation next step for every candidate.
  5. 5Take the strongest candidates to Compare and test exact model versions on the same tasks.

How to read the result

The qualitative label and requirement-match count apply to the model family, not individual versions; neither is a benchmark score. Verify each listed current version’s deployment availability and compare candidates on the same tasks.

What it cannot prove

The finder does not call models, verify current prices, or name a universal winner. Model families contain versions with different capabilities.

Common question: Why did a famous model not appear first?

The order reflects your selected constraints, not general popularity. Change one requirement to see what caused the result, then test exact versions.

Back to top
02

Compare specific model versions

Use real tasks and keep the decision project on this device.

Open this feature

What this feature is for

Compare two specific versions with a template or your own editable tasks.

Before you start

Open both provider services, confirm versions and settings, remove sensitive data, and remember that provider policies apply to prompts you submit there.

Step-by-step tutorial

  1. 1Choose two different specific model versions.
  2. 2Load a template or add and edit your own tasks.
  3. 3Run each prompt in both services and paste the complete answers.
  4. 4Review answers with model labels hidden; score five dimensions, choose Answer 1, Answer 2, or tie, and add notes.
  5. 5Reveal only after every required judgment, then download the V3 report or verify it locally.

How to read the result

The overall preference determines each task result. Close results or fewer than three completed tasks remain inconclusive.

What it cannot prove

Hiding labels only reduces visible-label bias. It is not double-blind, does not independently run models, and does not verify answer origin or establish a universal ranking.

Common question: What if scores and my preference disagree?

Keep the preference only when you can explain the difference in the note; otherwise revisit the scores.

Back to top
03

Verify evidence or a comparison report

Check supported JSON reports for structure, internal consistency, secret patterns, and a local fingerprint.

Open this feature

What this feature is for

Use the verifier to inspect an OpenGPT Evidence v1 or v2 package, or a V3 report downloaded from Decision Workspace. Every check runs locally in this browser.

Before you start

Obtain the original JSON file. For a comparison report, download it from Decision Workspace; complete every case first if you need a completed result rather than a draft.

Step-by-step tutorial

  1. 1Choose the JSON file; it stays in this browser.
  2. 2Confirm that the verifier recognizes Evidence v1/v2 or OpenGPT private comparison v1/v2 or Decision Workspace v3.
  3. 3For evidence packages, review required fields, runs, timestamps, scores, and secret-safety checks. For comparison reports, review candidate names, the task suite, draft or complete status, unique cases, and whether the summary matches the recorded verdicts.
  4. 4Read the declared summary and save the SHA-256 fingerprint if you need to confirm that the exact same file is reviewed later.
  5. 5If a check fails, correct the source data or finish the missing comparison cases instead of editing the downloaded report to force a pass.

How to read the result

For Evidence v1/v2, Passed means the package follows supported structure and consistency rules, while Demo marks an example. For a comparison report, Passed means a completed report is internally consistent; Draft means its structure is valid but the comparison is unfinished. Failed means at least one required check needs review.

What it cannot prove

A pass does not prove that the named model produced the pasted answers, that provider claims are true, or that the result generalizes beyond the recorded tests. Results from unrelated files are not automatically comparable.

Common question: Does a passed comparison report prove that one model is better?

No. It confirms only that the supported report is structurally valid and internally consistent. You must still judge the model identity, test quality, outputs, and decision relevance.

Back to top
04

Explore models

Understand model families before testing exact versions.

Open this feature

What this feature is for

Use the Models page to learn a family’s official positioning, common strengths, known trade-offs, and suitable work.

Before you start

Bring a task or constraint you care about. A family name alone is not enough for a purchasing or deployment decision.

Step-by-step tutorial

  1. 1Search by model, strength, limitation, or task.
  2. 2Read the visible family profiles, including strengths and limitations.
  3. 3Check suitable tasks and the linked official source.
  4. 4Select two exact versions from the expanded version lists.
  5. 5Use Compare versions to open the same-task comparison with both candidates.

How to read the result

A profile is orientation material. It helps form a testable hypothesis about fit; it is not an independent performance measurement.

What it cannot prove

Provider claims and family-level descriptions may not apply to every current or historical version. Confirm availability, price, privacy, and limits directly.

Common question: Can I choose a model from this page alone?

Use it to build a candidate list, then compare exact versions with your own tasks before a consequential decision.

Back to top
05

Understand the methodology

Learn what makes a model decision inspectable and reproducible.

Open this feature

What this feature is for

Use this page before designing, publishing, or relying on an evaluation. It separates evidence quality from model popularity.

Before you start

Start with a concrete decision question, named model versions, representative tasks, and a plan for preserving raw outputs.

Step-by-step tutorial

  1. 1Write the decision and success criteria before seeing results.
  2. 2Use the same task set, settings, and scoring rules for every candidate.
  3. 3Preserve prompts, outputs, settings, failures, timestamps, and suite fingerprints.
  4. 4State what was manual, automated, reviewed with model labels hidden, randomized, or independently reproduced.
  5. 5Publish conclusions together with uncertainty and explicit limitations.

How to read the result

Evidence v1/v2 mainly checks format and declarations. Decision Workspace V3 supports local recomputation (L3); L4 independent reproduction remains unsupported.

What it cannot prove

A rigorous process can reduce bias but cannot guarantee that a sample represents every user, language, future version, or production condition.

Common question: Is a higher evidence level always necessary?

Match rigor to consequence. A personal trial may need L1; a public or high-stakes claim should seek stronger independent evidence.

Back to top

Plain-language glossary

Use these definitions when a label is unfamiliar.

Requirement fit
A qualitative family-fit label plus the number of matched requirements; neither rates individual versions. Verify each listed current version’s deployment availability and compare it on the same tasks.
Model family
A group of related model versions. Capabilities can differ across versions.
Benchmark
A repeatable task set and scoring rule used to compare systems under stated conditions.
Evidence file
A structured record that declares how an evaluation was run and what it observed.
Schema
The machine-readable rules that define required fields and allowed values in a file.
SHA-256 fingerprint
A local identifier for exact file bytes; changing the file changes the fingerprint.

Troubleshooting

No model meets every requirement

Review which requirement excluded each candidate, then relax only a constraint that is genuinely flexible.

Review requirements

An exact version is missing

Check the Models page for listed versions and official sources. Do not substitute a family name for an exact version.

Check model versions

The comparison report is incomplete

Return to every marked case and provide both complete outputs plus a verdict.

Continue comparison

A local draft conflicts or needs a reset

Return to Compare, review the newer local draft warning, then keep or reset the draft intentionally.

Review local draft

An evidence file fails

Use the failed checklist item to correct the source record. Do not insert secrets or rewrite outcomes.

Open verifier
contact@opengpt.com

Still need help?

Report a broken link, incorrect model detail, or issue you could not resolve.

Email support