GUIDES & SUPPORT

OpenGPT Help Center

Find the right workflow, follow each step, and understand what the result can—and cannot—tell you.

Start with what you want to do

Each workflow answers a different question. Start with the closest goal, then follow the linked guide.

1Model rankingsUnderstand published signals and identity status.2Agent rankingsChoose complete agent configurations by task and evidence.3ChooseNarrow candidates from real requirements.4Compare public dataPlace two to four catalog model IDs side by side.5Save modelsKeep a local shortlist in this browser.6Evaluate real answersTest two exact versions on the same tasks.7Submit or correct evidenceSend a public model or Agent source for editorial review.8Verify a reportCheck an OpenGPT JSON report locally before sharing it.

01

AI model rankings

See what is ranked, where the data came from, and which records include an official model ID.

Open

What this feature is for

Use each ranking angle to inspect published signals, review source records, and build a shortlist for real-task testing.

Before you start

Start with the question you need to answer. User preference, overall capability, objective tasks, cost, and open-weight rankings use different sources and scales.

Step-by-step tutorial

  1. 1Choose the ranking angle that matches your decision.
  2. 2Use Full source ranking for every publisher record, or Official model IDs for records linked by provider documentation.
  3. 3Read source position, score range, votes, license, and configuration together.
  4. 4Open Model details on any row. Exact official mappings use the official profile; every other row opens a source profile that preserves the published name and configuration.
  5. 5Add any row to comparison. Exact, configuration-safe mappings use official-version comparison; all other rows use source-record comparison within the same source, dimension, and snapshot.

How to read the result

A rank is relative to one source and scope. Ties can make rank numbers skip, and scores from different ranking angles are not directly comparable.

What it cannot prove

OpenGPT organizes third-party public data and did not rerun these evaluations. No ranking identifies the best model for every user or task.

Common question: Why do some rows show only a source model ID or name?

Every row now has model details and a comparison action. OpenGPT adds an official model ID only when provider documentation proves the exact mapping; otherwise source-record comparison preserves the publisher name, configuration, and ranking data without inventing a new official ID.

Back to top
02

AI agent rankings

Choose an agent by task, complete configuration, evidence quality, and comparable results.

Open

What this feature is for

Explore 96 published Agent evaluation configurations across seven task scopes, plus 109 separate BFCL tool-calling model results, without turning incompatible evidence into one score.

Before you start

An agent result belongs to one complete configuration: base model, scaffold, tools, permissions, budget, environment, benchmark version, and date. OpenGPT preserves fields reported by the source and leaves missing fields unreported—not zero.

Step-by-step tutorial

  1. 1Choose the task category that matches your work.
  2. 2Check the benchmark scope, version, evaluation date, freshness, and audit note before reading the rank.
  3. 3Use the evidence grade to distinguish reproduced or checked results from transparent public runs and vendor-reported results.
  4. 4Compare success, repeat-run evidence, duration, and intervention only when the source reports them under the same conditions. OpenGPT shows cost per success only when the published total and task/run denominator use the same verified scope.
  5. 5Follow useful configurations, review their history, and select two to four entries from the same comparable group for a side-by-side view.

How to read the result

A rank describes one complete system configuration within one benchmark scope. Read it with the evidence grade, freshness, sample or run count, and any reported cost and reliability fields.

What it cannot prove

OpenGPT does not create a universal Agent score. Unless a result is explicitly marked as reproduced, OpenGPT is organizing upstream evidence rather than rerunning it; no benchmark proves production readiness for every workflow.

Common question: Why can the same agent or model appear more than once?

A different scaffold, model version, tool set, permission level, budget, or benchmark setting creates a different agent configuration and must remain a separate record.

Back to top
03

Submit a model or Agent

Send a public identity or evidence record for review without changing the rankings.

Open

What this feature is for

Use the submission flow when an official model ID or Agent product is missing, or when a public source needs editorial review.

Before you start

Sign in with a verified account and prepare a public HTTPS source. Never submit credentials, private files, or personal information.

Step-by-step tutorial

  1. 1Choose Model or Agent at the top of the form.
  2. 2Enter the exact public name, provider or organization, and the requested identity fields.
  3. 3Choose the source type and link to official documentation, a public repository, or a public benchmark result.
  4. 4Explain what the reviewer should verify, then submit once.
  5. 5Track the review status under your recent submissions. An approved Agent product is automatically listed in the searchable Agent directory; ranking remains separate.

How to read the result

A successful submission receives a reference number and enters the review queue. An approved Agent product is a public directory record, not a leaderboard rank or OpenGPT endorsement.

What it cannot prove

Submitters cannot set rank, evidence grade, comparability, or publication status. Duplicate and rate limits reduce abuse, and approval never changes a ranking automatically.

Common question: Why was an approved submission not added to the ranking?

Approved Agent products are searchable in the directory, but a ranking still requires comparable public evaluation evidence.

Back to top
04

Leaderboard changes and version history

Track movement across comparable snapshots and compare catalog model IDs.

Open

What this feature is for

Use this page to see what moved since the prior source snapshot, which records first appeared, how exact versions differ, and when data was retrieved.

Before you start

Historical movement currently uses the LMArena user-preference scope. Other sources stay current-only until a comparable prior snapshot exists.

Step-by-step tutorial

  1. 1Check the source date, comparison baseline, retrieval date, and comparison rule.
  2. 2Read Fastest risers only among records present in both snapshots.
  3. 3Treat First appeared as a new source entry—not an official release date.
  4. 4Filter the movement table by up, down, unchanged, first appeared, or no longer listed.
  5. 5Compare two catalog model IDs, then download a source snapshot when you need an audit trail.

How to read the result

OpenGPT calculates movement only when source, scope, unit, and method match, and keeps downloadable monthly snapshots.

What it cannot prove

Only release dates documented by the provider qualify as product releases. Version comparison does not infer speed, price, context length, or quality from names.

Common question: Why do Artificial Analysis and LiveBench show no rank movement?

OpenGPT has not frozen a comparable prior edition for those sources. Missing history is shown instead of creating a synthetic change.

Back to top
05

Rankings by task

Start from a real work scenario and review three official model versions worth testing first.

Open

What this feature is for

Use task rankings to turn broad public signals into a small, practical shortlist for coding, research, documents, support, multilingual work, multimodal work, cost, privacy, or speed.

Before you start

Choose the task closest to your work, then note your actual prompts, data, tools, region, budget, and privacy requirements. The public evidence shown may cover a broader scope than your task.

Step-by-step tutorial

  1. 1Open the task scenario closest to the decision you need to make.
  2. 2Read why the three official versions were selected and what remains uncertain.
  3. 3Inspect the public source, source rank or value, mapping review date, and data cutoff for each candidate.
  4. 4Select two or three candidates for a side-by-side public-data comparison.
  5. 5Open the real-task evaluation with the first two candidates, edit the sample task set, and test both with comparable settings.

How to read the result

You get three official starting candidates, visible public evidence where available, and direct paths to compare or run a real-task test.

What it cannot prove

A task ranking is a starting shortlist, not a universal ranking or proof of the best model. Missing evidence is not zero, and unlike public sources are not combined into one score.

Common question: Why can a starting candidate have no comparable public result?

An official version may be relevant to the task by product position or deployment type even when no exact-version result is mapped in the selected public source. The missing evidence stays visible and must be resolved through your own test.

Back to top
06

Model Upgrade Radar

Set the models you use, recurring work, and priorities, then review cautious signals for when a fair comparison is worthwhile.

Open

What this feature is for

Use the radar to answer a practical question: is there enough comparable public evidence to test an alternative to the model you use now?

Before you start

Choose exact official versions, at least one recurring task, and up to two priorities. Use the same browser and device because the setup stays local.

Step-by-step tutorial

  1. 1Choose the exact official versions you currently use; saved models are not assumed to be current.
  2. 2Select recurring tasks and up to two priorities such as quality, cost, speed, privacy, or long context.
  3. 3Review whether each model is worth comparing, safe to keep watching, or approaching retirement.
  4. 4Inspect comparable public evidence, then test both models on the same real work.
  5. 5Review weekly changes or subscribe to the public ranking-changes RSS feed.

How to read the result

The radar creates a device-local upgrade check with a current model, task-matched alternative when available, evidence date, uncertainty boundary, and direct comparison actions.

What it cannot prove

An upgrade check is not an instruction to switch. Cost, speed, privacy, or context may still require direct testing, and local settings are not synced or backed up.

Common question: Does “worth comparing” mean I should replace my current model?

No. It means a comparable signal or lifecycle change makes a same-task test worthwhile. Switch only after the alternative succeeds on your real work and constraints.

Back to top
07

Choose an AI model

Turn four practical requirements into an explainable shortlist.

Open

What this feature is for

Use the finder when you know the work you need to do but do not know which model series to investigate first.

Before you start

Know the main task, where the model may run, your most important operating priority, and any language requirement.

Step-by-step tutorial

  1. 1Choose the task closest to your real work.
  2. 2Set the deployment boundary, especially when data cannot leave your environment.
  3. 3Choose the priority you are least willing to compromise.
  4. 4Set the language need, then review the match reasons, risks, and validation next step for every candidate.
  5. 5Take the strongest candidates to real-task evaluation and test exact model versions on the same tasks.

How to read the result

The qualitative label and requirement-match count apply to the model series, not individual versions; neither is a benchmark score. Verify each listed current version’s deployment availability and compare candidates on the same tasks.

What it cannot prove

The finder does not call models, verify current prices, or name a universal winner. A model series can include versions with different capabilities.

Common question: Why did a famous model not appear first?

The order reflects your selected constraints, not general popularity. Change one requirement to see what caused the result, then test exact versions.

Back to top
08

Model directory and details

Browse provider-documented model IDs, open dedicated detail pages, and check whether comparable public ranking data exists.

Open

What this feature is for

Use the directory when you need an exact official model ID, provider destination, suitable and unsuitable tasks, sources, and a clear ranking-data status.

Before you start

Bring a model name, provider, deployment need, or task. A provider-documented model ID does not guarantee comparable public evaluation data.

Step-by-step tutorial

  1. 1Search by model, version, provider, or task.
  2. 2Filter by deployment method and by ranked or currently unranked status.
  3. 3Open a model detail page to review identity, intended use, public-data coverage, strengths, limitations, sources, and review date.
  4. 4Use the official destination to verify current availability, pricing, privacy, and limits.
  5. 5Add two to four catalog versions to data comparison, then take two finalists to real-task evaluation.

How to read the result

Ranked means comparable published data is linked to that exact model ID. Unranked means this edition has no suitable comparable record; it is not a zero score.

What it cannot prove

Provider documentation establishes the catalog identity, not current availability in every country or service tier. Ranking coverage, model details, and provider terms can change after the review date.

Common question: Does unranked mean the model is weak?

No. It means this edition lacks suitable comparable public data for that model ID. Test it directly if it remains relevant to your decision.

Back to top
09

Data comparison

Compare two to four exact official versions, or records from one published ranking view.

Open

What this feature is for

Official-version comparison organizes provider-documented IDs across published sources. Source-record comparison preserves the publisher’s name and settings when an exact official comparison would lose information.

Before you start

For official versions, choose two to four catalog IDs. For source records, choose two to four rows from the same source, dimension, and snapshot.

Step-by-step tutorial

  1. 1Add any row from a ranking. Configuration-safe mappings use official-version comparison; all other rows use source-record comparison.
  2. 2Official results remain separated by source and dimension.
  3. 3Source-record results preserve the published name, settings, rank, and score on one common source scale.
  4. 4Missing data remains unknown rather than zero.
  5. 5Only exact official versions can continue to real-task evaluation after you choose two finalists.

How to read the result

The comparison creates an evidence-aware shortlist and shows where information is missing. It does not produce a new benchmark score or declare a universal winner.

What it cannot prove

OpenGPT does not call the models in this view. Published scores from different sources or scopes are not interchangeable, and model details can change after their review date.

Common question: How is data comparison different from real-task evaluation?

Data comparison summarizes already published information for two to four models. Real-task evaluation takes exactly two versions, asks you to run the same prompts in both services, and reviews the pasted answers with visible model labels hidden.

Back to top
10

Saved models

Keep a small working list of ranking entries in this browser.

Open

What this feature is for

Use Saved models to return quickly to candidates, open trusted destinations, and move supported model IDs into a comparison.

Before you start

Use the same browser and device. The list is stored locally, without an OpenGPT account or cloud synchronization.

Step-by-step tutorial

  1. 1Select Save on any ranking row.
  2. 2Open Saved models from the top navigation to review saved entries.
  3. 3Every saved source record shows its source model ID or source name. Records with an official model ID can also open the provider destination or enter comparison.
  4. 4Add two different supported model IDs, then open the public-data comparison. Choose two finalists there before starting a real-task evaluation.
  5. 5Remove entries you no longer need. Refreshing or changing the site language keeps the list in the same browser.

How to read the result

The count beside Saved models shows how many entries are saved. Source identity and official model identity remain separate.

What it cannot prove

The list is not an account, bookmark sync, or backup. Another browser or device will not see it, and private browsing, storage restrictions, or clearing site data can remove it.

Common question: Why did my saved models disappear?

Check that you are using the same browser profile and device. If site storage was cleared or blocked, OpenGPT cannot restore the local list.

Back to top
11

Real-task evaluation

Use real tasks and keep the decision project on this device.

Open

What this feature is for

Compare two specific versions with a template or your own editable tasks.

Before you start

Choose two different catalog model IDs that support the same-task workflow. Open both provider services, check versions and settings, remove sensitive data, and remember that provider policies apply to prompts submitted there.

Step-by-step tutorial

  1. 1Choose two supported specific versions from the directory, ranking, or My Models. The same official version cannot fill both positions.
  2. 2Load a template or add and edit your own tasks.
  3. 3Run each prompt in both services and paste the complete answers.
  4. 4Review answers with model labels hidden; score five dimensions, choose Answer 1, Answer 2, or tie, and add notes.
  5. 5Reveal only after every required judgment, then download the V3 report or verify it locally.

How to read the result

The overall preference determines each task result. Close results or fewer than three completed tasks remain inconclusive.

What it cannot prove

Hiding labels only reduces visible-label bias. It is not double-blind, does not independently run models, and does not verify answer origin or establish a universal ranking.

Common question: Why can’t I add some records to data comparison?

A source record without a provider-documented model ID cannot enter the catalog comparison. The same model may also already be selected, or the four-model list may be full.

Back to top
12

Saved task sets and re-evaluation

Reuse the same real prompts in a later model evaluation without rebuilding the task set.

Open

What this feature is for

Use saved task sets to make future comparisons more consistent and to re-evaluate exact model versions after a release or public-data change.

Before you start

Finish the task titles and prompts, choose a clear task-set name, and use the same browser and device. A task set can contain sensitive work text, so remove private data before saving.

Step-by-step tutorial

  1. 1In Real-task evaluation, prepare or edit a task set that represents your work.
  2. 2Give the task set a recognizable name and save it.
  3. 3Later, choose the saved task set and load it into the current evaluation.
  4. 4Select two exact model versions and run every prompt with comparable settings.
  5. 5From a completed result, choose Re-evaluate to keep the tasks and model versions while clearing old answers, judgments, and the prior report.

How to read the result

The task-set name, template, language, task titles, and prompts are stored locally for reuse. Re-evaluation starts a clean run while preserving the intended test inputs.

What it cannot prove

Saved task sets exist only on this device and are not synced or backed up. Answers, model choices, reviews, reports, and API keys are not saved in a task set. Clearing site data can remove all saved sets.

Common question: What changes when I choose Re-evaluate?

OpenGPT keeps the same tasks and exact model versions, then clears prior answers, scores, judgments, reveal state, and report so the next run is independent of the old result.

Back to top
13

Shareable decision report

Create a privacy-safe link for an official-model shortlist and its public-data cutoff.

Open

What this feature is for

Use the report to share which task and exact official versions are under consideration, the public evidence available, and an optional selected candidate.

Before you start

Create the report from a task ranking. Confirm the task, exact model IDs, data cutoff, and optional choice before sharing.

Step-by-step tutorial

  1. 1Open a task ranking and create a report from two or three official candidates.
  2. 2Optionally mark one model as selected; leaving the choice blank is valid.
  3. 3Review public-data coverage and the report boundary.
  4. 4Copy the link, use the device share sheet, print it, or download the sanitized JSON record.
  5. 5A recipient can open the same shortlist, compare public data, or start a real-task evaluation.

How to read the result

The shared URL carries only the task ID, exact official model IDs, public-data cutoff, and optional selected model, so the recipient can inspect the same shortlist.

What it cannot prove

The shareable report contains no prompts, model answers, notes, API keys, or full evaluation report. It is a shortlist record, not proof of a universal winner or authenticated provider output.

Common question: Can someone see my evaluation content from the share link?

No. Prompts, answers, notes, API keys, and the local evaluation project are excluded. Always inspect the link before sharing if you added information outside the supported report flow.

Back to top
14

Report verification

Check supported JSON reports for structure, internal consistency, secret patterns, and a local fingerprint.

Open

What this feature is for

Use the verifier to inspect an OpenGPT Evidence v1 or v2 package, or a V3 report downloaded from Decision Workspace. Every check runs locally in this browser.

Before you start

Obtain the original JSON file. For a comparison report, download it from Decision Workspace; complete every case first if you need a completed result rather than a draft.

Step-by-step tutorial

  1. 1Choose the JSON file; it stays in this browser.
  2. 2Confirm that the verifier recognizes Evidence v1/v2 or OpenGPT private comparison v1/v2 or Decision Workspace v3.
  3. 3For evidence packages, review required fields, runs, timestamps, scores, and secret-safety checks. For comparison reports, review candidate names, the task suite, draft or complete status, unique cases, and whether the summary matches the recorded verdicts.
  4. 4Read the declared summary and save the SHA-256 fingerprint if you need to confirm that the exact same file is reviewed later.
  5. 5If a check fails, correct the source data or finish the missing comparison cases instead of editing the downloaded report to force a pass.

How to read the result

For Evidence v1/v2, Passed means the package follows supported structure and consistency rules, while Demo marks an example. For a comparison report, Passed means a completed report is internally consistent; Draft means its structure is valid but the comparison is unfinished. Failed means at least one required check needs review.

What it cannot prove

A pass does not prove that the named model produced the pasted answers, that provider claims are true, or that the result generalizes beyond the recorded tests. Results from unrelated files are not automatically comparable.

Common question: Does a passed comparison report prove that one model is better?

No. It confirms only that the supported report is structurally valid and internally consistent. You must still judge the model identity, test quality, outputs, and decision relevance.

Back to top
15

Evaluation methodology

Learn what makes a model decision inspectable and reproducible.

Open

What this feature is for

Use this page before designing, publishing, or relying on an evaluation. It separates evidence quality from model popularity.

Before you start

Start with a concrete decision question, named model versions, representative tasks, and a plan for preserving raw outputs.

Step-by-step tutorial

  1. 1Write the decision and success criteria before seeing results.
  2. 2Use the same task set, settings, and scoring rules for every candidate.
  3. 3Preserve prompts, outputs, settings, failures, timestamps, and suite fingerprints.
  4. 4State what was manual, automated, reviewed with model labels hidden, randomized, or independently reproduced.
  5. 5Publish conclusions together with uncertainty and explicit limitations.

How to read the result

Evidence v1/v2 mainly checks format and declarations. Decision Workspace V3 supports local recomputation (L3); L4 independent reproduction remains unsupported.

What it cannot prove

A rigorous process can reduce bias but cannot guarantee that a sample represents every user, language, future version, or production condition.

Common question: Is a higher evidence level always necessary?

Match rigor to consequence. A personal trial may need L1; a public or high-stakes claim should seek stronger independent evidence.

Back to top

Plain-language glossary

Use these definitions when a label is unfamiliar.

Source record
A row preserved from a published ranking, with the source model ID or name kept exactly for review and audit.
Official model ID
A provider-documented model identifier that OpenGPT keeps separate from the source record name.
Catalog version
A model version listed from provider documentation, with its official destination and review source.
Likely score range
A published uncertainty interval around the displayed score. Wider ranges mean less precision; ranges from different ranking sources are not directly comparable.
My Models
A shortlist stored only in this browser. It is not an account, cloud sync, or provider bookmark.
Requirement fit
A qualitative family-fit label plus the number of matched requirements; neither rates individual versions. Verify availability and compare exact versions on the same tasks.
Model family
A group of related model versions. Capabilities, prices, and access can differ across versions.
Benchmark
A repeatable task set and scoring rule used to compare systems under stated conditions.
Agent configuration
The base model, scaffold, tools, permissions, budget, environment, benchmark version, and date that produced one Agent result.
Evidence grade
A label describing how a result was produced or checked. It does not change the source score.
Cost per success
Reported evaluation cost divided by successful tasks. It is shown only when comparable source data exists; missing cost is not zero.
Evidence file
A structured record that declares how an evaluation was run and what it observed.
Schema
The machine-readable rules that define required fields and allowed values in a file.
SHA-256 fingerprint
A local identifier for exact file bytes; changing the file changes the fingerprint.

Troubleshooting

No model meets every requirement

Review which requirement excluded each candidate, then relax only a constraint that is genuinely flexible.

Review requirements

A source rank number skips

The publisher may use tied ranks or multiple configurations. OpenGPT preserves the original position instead of renumbering the visible rows.

Review the ranking

An Agent result has missing cost or reliability

The source did not publish a comparable value. OpenGPT leaves it unreported instead of treating it as zero; compare only values from the same benchmark scope.

Review Agent evidence

A row has no official model ID

It is a complete source record rather than an unresolved status. Open the original source; use the model directory when you need provider-documented identity and official access.

Review source identity

My Models is empty or disappeared

Use the same browser profile and device. Private mode, blocked storage, or clearing site data can remove the local list.

Open My Models

A record cannot be added to data comparison

Data comparison accepts catalog model IDs. The same model may already be selected, or the list may have reached its four-model limit. Choose another catalog version or remove a candidate first.

Open the model directory

The comparison report is incomplete

Return to every marked case and provide both complete outputs plus a verdict.

Continue comparison

A local draft conflicts or needs a reset

Return to Compare, review the newer local draft warning, then keep or reset the draft intentionally.

Review local draft

An evidence file fails

Use the failed checklist item to correct the source record. Do not insert secrets or rewrite outcomes.

Open verifier
contact@opengpt.com

Still need help?

Report a broken link, incorrect model detail, or issue you could not resolve.

Email support