GUIDES & SUPPORT

OpenGPT Help Center

Choose a guide, follow the steps, or search for a problem.

Start here: three different jobs

Choosing, comparing, and verifying support different decisions. None of them is a universal leaderboard.

1ChooseNarrow model families from your requirements.2CompareCompare real outputs from exact versions on the same tasks.3VerifyCheck the structure and internal consistency of an evidence file.

01

Choose an AI model

Turn four practical requirements into an explainable shortlist.

Open this feature

What this feature is for

Use the finder when you know the work you need to do but do not know which model family to investigate first.

Before you start

Know the main task, where the model may run, your most important operating priority, and any language requirement.

Step-by-step tutorial

  1. 1Choose the task closest to your real work.
  2. 2Set the deployment boundary, especially when data cannot leave your environment.
  3. 3Choose the priority you are least willing to compromise.
  4. 4Set the language need, then review the match reasons, risks, and validation next step for every candidate.
  5. 5Take the strongest candidates to Compare and test exact model versions on the same tasks.

How to read the result

The qualitative label and requirement-match count apply to the model family, not individual versions; neither is a benchmark score. Verify each listed current version’s deployment availability and compare candidates on the same tasks.

What it cannot prove

The finder does not call models, verify current prices, or name a universal winner. Model families contain versions with different capabilities.

Common question: Why did a famous model not appear first?

The order reflects your selected constraints, not general popularity. Change one requirement to see what caused the result, then test exact versions.

Back to top
02

Compare two models privately

Record a fair same-prompt comparison without giving OpenGPT an API key.

Open this feature

What this feature is for

Use Compare to test two exact model versions on one of three built-in practical task sets.

Before you start

Open both model services yourself. Confirm their exact versions and settings, remove sensitive data, and remember that those services receive your prompts and may charge for access.

Step-by-step tutorial

  1. 1Select two different exact model versions.
  2. 2Choose one of the three built-in task sets.
  3. 3Use each provider link and run the same prompt in both services without changing it.
  4. 4Paste both complete outputs, then judge accuracy, instruction following, completeness, safety, and clarity.
  5. 5Finish all cases, review incomplete items, and download the local report.

How to read the result

Wins and ties summarize only the completed cases in this session. Notes are important evidence for why one answer was preferred.

What it cannot prove

OpenGPT does not call the models, verify response origin, blind candidate labels, or accept custom prompts in the current workflow. A small manual comparison is not a universal leaderboard.

Common question: What should I choose when both answers have problems?

Choose a tie when neither is clearly better, and record the problems in the note. Do not force a winner when evidence is weak.

Back to top
03

Verify evidence or a comparison report

Check supported JSON reports for structure, internal consistency, secret patterns, and a local fingerprint.

Open this feature

What this feature is for

Use the verifier to inspect an OpenGPT Evidence v1 or v2 package, or a report downloaded from Private comparison. Every check runs locally in this browser.

Before you start

Obtain the original JSON file. For a comparison report, download it from Private comparison; complete every case first if you need a completed result rather than a draft.

Step-by-step tutorial

  1. 1Choose the JSON file; it stays in this browser.
  2. 2Confirm that the verifier recognizes Evidence v1/v2 or OpenGPT private comparison v1/v2.
  3. 3For evidence packages, review required fields, runs, timestamps, scores, and secret-safety checks. For comparison reports, review candidate names, the task suite, draft or complete status, unique cases, and whether the summary matches the recorded verdicts.
  4. 4Read the declared summary and save the SHA-256 fingerprint if you need to confirm that the exact same file is reviewed later.
  5. 5If a check fails, correct the source data or finish the missing comparison cases instead of editing the downloaded report to force a pass.

How to read the result

For Evidence v1/v2, Passed means the package follows supported structure and consistency rules, while Demo marks an example. For a comparison report, Passed means a completed report is internally consistent; Draft means its structure is valid but the comparison is unfinished. Failed means at least one required check needs review.

What it cannot prove

A pass does not prove that the named model produced the pasted answers, that provider claims are true, or that the result generalizes beyond the recorded tests. Results from unrelated files are not automatically comparable.

Common question: Does a passed comparison report prove that one model is better?

No. It confirms only that the supported report is structurally valid and internally consistent. You must still judge the model identity, test quality, outputs, and decision relevance.

Back to top
04

Explore models

Understand model families before testing exact versions.

Open this feature

What this feature is for

Use the Models page to learn a family’s official positioning, common strengths, known trade-offs, and suitable work.

Before you start

Bring a task or constraint you care about. A family name alone is not enough for a purchasing or deployment decision.

Step-by-step tutorial

  1. 1Search by model, strength, limitation, or task.
  2. 2Read the visible family profiles, including strengths and limitations.
  3. 3Check suitable tasks and the linked official source.
  4. 4Select two exact versions from the expanded version lists.
  5. 5Use Compare versions to open the same-task comparison with both candidates.

How to read the result

A profile is orientation material. It helps form a testable hypothesis about fit; it is not an independent performance measurement.

What it cannot prove

Provider claims and family-level descriptions may not apply to every current or historical version. Confirm availability, price, privacy, and limits directly.

Common question: Can I choose a model from this page alone?

Use it to build a candidate list, then compare exact versions with your own tasks before a consequential decision.

Back to top
05

Understand the methodology

Learn what makes a model decision inspectable and reproducible.

Open this feature

What this feature is for

Use this page before designing, publishing, or relying on an evaluation. It separates evidence quality from model popularity.

Before you start

Start with a concrete decision question, named model versions, representative tasks, and a plan for preserving raw outputs.

Step-by-step tutorial

  1. 1Write the decision and success criteria before seeing results.
  2. 2Use the same task set, settings, and scoring rules for every candidate.
  3. 3Preserve prompts, outputs, settings, failures, timestamps, and suite fingerprints.
  4. 4State what was manual, automated, blinded, randomized, or independently reproduced.
  5. 5Publish conclusions together with uncertainty and explicit limitations.

How to read the result

Evidence levels describe a target standard. L1 is currently supported, L2 is partial, and L3–L4 require capabilities outside the current local workflow.

What it cannot prove

A rigorous process can reduce bias but cannot guarantee that a sample represents every user, language, future version, or production condition.

Common question: Is a higher evidence level always necessary?

Match rigor to consequence. A personal trial may need L1; a public or high-stakes claim should seek stronger independent evidence.

Back to top

Plain-language glossary

Use these definitions when a label is unfamiliar.

Requirement fit
A qualitative family-fit label plus the number of matched requirements; neither rates individual versions. Verify each listed current version’s deployment availability and compare it on the same tasks.
Model family
A group of related model versions. Capabilities can differ across versions.
Benchmark
A repeatable task set and scoring rule used to compare systems under stated conditions.
Evidence file
A structured record that declares how an evaluation was run and what it observed.
Schema
The machine-readable rules that define required fields and allowed values in a file.
SHA-256 fingerprint
A local identifier for exact file bytes; changing the file changes the fingerprint.

Troubleshooting

No model meets every requirement

Review which requirement excluded each candidate, then relax only a constraint that is genuinely flexible.

Review requirements

An exact version is missing

Check the Models page for listed versions and official sources. Do not substitute a family name for an exact version.

Check model versions

The comparison report is incomplete

Return to every marked case and provide both complete outputs plus a verdict.

Continue comparison

A local draft conflicts or needs a reset

Return to Compare, review the newer local draft warning, then keep or reset the draft intentionally.

Review local draft

An evidence file fails

Use the failed checklist item to correct the source record. Do not insert secrets or rewrite outcomes.

Open verifier
contact@opengpt.com

Still need help?

Report a broken link, incorrect model detail, or issue you could not resolve.

Email support