No model meets every requirement
Review which requirement excluded each candidate, then relax only a constraint that is genuinely flexible.
Review requirementsFind the right workflow, follow each step, and understand what the result can—and cannot—tell you.
Each workflow answers a different question. Start with the closest goal, then follow the linked guide.
See what is ranked, where the data came from, and which records include an official model ID.
Use each ranking angle to inspect published signals, review source records, and build a shortlist for real-task testing.
Start with the question you need to answer. User preference, overall capability, objective tasks, cost, and open-weight rankings use different sources and scales.
A rank is relative to one source and scope. Ties can make rank numbers skip, and scores from different ranking angles are not directly comparable.
OpenGPT organizes third-party public data and did not rerun these evaluations. No ranking identifies the best model for every user or task.
Every row now has model details and a comparison action. OpenGPT adds an official model ID only when provider documentation proves the exact mapping; otherwise source-record comparison preserves the publisher name, configuration, and ranking data without inventing a new official ID.
Choose an agent by task, complete configuration, evidence quality, and comparable results.
Explore 96 published Agent evaluation configurations across seven task scopes, plus 109 separate BFCL tool-calling model results, without turning incompatible evidence into one score.
An agent result belongs to one complete configuration: base model, scaffold, tools, permissions, budget, environment, benchmark version, and date. OpenGPT preserves fields reported by the source and leaves missing fields unreported—not zero.
A rank describes one complete system configuration within one benchmark scope. Read it with the evidence grade, freshness, sample or run count, and any reported cost and reliability fields.
OpenGPT does not create a universal Agent score. Unless a result is explicitly marked as reproduced, OpenGPT is organizing upstream evidence rather than rerunning it; no benchmark proves production readiness for every workflow.
A different scaffold, model version, tool set, permission level, budget, or benchmark setting creates a different agent configuration and must remain a separate record.
Send a public identity or evidence record for review without changing the rankings.
Use the submission flow when an official model ID or Agent product is missing, or when a public source needs editorial review.
Sign in with a verified account and prepare a public HTTPS source. Never submit credentials, private files, or personal information.
A successful submission receives a reference number and enters the review queue. An approved Agent product is a public directory record, not a leaderboard rank or OpenGPT endorsement.
Submitters cannot set rank, evidence grade, comparability, or publication status. Duplicate and rate limits reduce abuse, and approval never changes a ranking automatically.
Approved Agent products are searchable in the directory, but a ranking still requires comparable public evaluation evidence.
Track movement across comparable snapshots and compare catalog model IDs.
Use this page to see what moved since the prior source snapshot, which records first appeared, how exact versions differ, and when data was retrieved.
Historical movement currently uses the LMArena user-preference scope. Other sources stay current-only until a comparable prior snapshot exists.
OpenGPT calculates movement only when source, scope, unit, and method match, and keeps downloadable monthly snapshots.
Only release dates documented by the provider qualify as product releases. Version comparison does not infer speed, price, context length, or quality from names.
OpenGPT has not frozen a comparable prior edition for those sources. Missing history is shown instead of creating a synthetic change.
Start from a real work scenario and review three official model versions worth testing first.
Use task rankings to turn broad public signals into a small, practical shortlist for coding, research, documents, support, multilingual work, multimodal work, cost, privacy, or speed.
Choose the task closest to your work, then note your actual prompts, data, tools, region, budget, and privacy requirements. The public evidence shown may cover a broader scope than your task.
You get three official starting candidates, visible public evidence where available, and direct paths to compare or run a real-task test.
A task ranking is a starting shortlist, not a universal ranking or proof of the best model. Missing evidence is not zero, and unlike public sources are not combined into one score.
An official version may be relevant to the task by product position or deployment type even when no exact-version result is mapped in the selected public source. The missing evidence stays visible and must be resolved through your own test.
Set the models you use, recurring work, and priorities, then review cautious signals for when a fair comparison is worthwhile.
Use the radar to answer a practical question: is there enough comparable public evidence to test an alternative to the model you use now?
Choose exact official versions, at least one recurring task, and up to two priorities. Use the same browser and device because the setup stays local.
The radar creates a device-local upgrade check with a current model, task-matched alternative when available, evidence date, uncertainty boundary, and direct comparison actions.
An upgrade check is not an instruction to switch. Cost, speed, privacy, or context may still require direct testing, and local settings are not synced or backed up.
No. It means a comparable signal or lifecycle change makes a same-task test worthwhile. Switch only after the alternative succeeds on your real work and constraints.
Turn four practical requirements into an explainable shortlist.
Use the finder when you know the work you need to do but do not know which model series to investigate first.
Know the main task, where the model may run, your most important operating priority, and any language requirement.
The qualitative label and requirement-match count apply to the model series, not individual versions; neither is a benchmark score. Verify each listed current version’s deployment availability and compare candidates on the same tasks.
The finder does not call models, verify current prices, or name a universal winner. A model series can include versions with different capabilities.
The order reflects your selected constraints, not general popularity. Change one requirement to see what caused the result, then test exact versions.
Browse provider-documented model IDs, open dedicated detail pages, and check whether comparable public ranking data exists.
Use the directory when you need an exact official model ID, provider destination, suitable and unsuitable tasks, sources, and a clear ranking-data status.
Bring a model name, provider, deployment need, or task. A provider-documented model ID does not guarantee comparable public evaluation data.
Ranked means comparable published data is linked to that exact model ID. Unranked means this edition has no suitable comparable record; it is not a zero score.
Provider documentation establishes the catalog identity, not current availability in every country or service tier. Ranking coverage, model details, and provider terms can change after the review date.
No. It means this edition lacks suitable comparable public data for that model ID. Test it directly if it remains relevant to your decision.
Compare two to four exact official versions, or records from one published ranking view.
Official-version comparison organizes provider-documented IDs across published sources. Source-record comparison preserves the publisher’s name and settings when an exact official comparison would lose information.
For official versions, choose two to four catalog IDs. For source records, choose two to four rows from the same source, dimension, and snapshot.
The comparison creates an evidence-aware shortlist and shows where information is missing. It does not produce a new benchmark score or declare a universal winner.
OpenGPT does not call the models in this view. Published scores from different sources or scopes are not interchangeable, and model details can change after their review date.
Data comparison summarizes already published information for two to four models. Real-task evaluation takes exactly two versions, asks you to run the same prompts in both services, and reviews the pasted answers with visible model labels hidden.
Keep a small working list of ranking entries in this browser.
Use Saved models to return quickly to candidates, open trusted destinations, and move supported model IDs into a comparison.
Use the same browser and device. The list is stored locally, without an OpenGPT account or cloud synchronization.
The count beside Saved models shows how many entries are saved. Source identity and official model identity remain separate.
The list is not an account, bookmark sync, or backup. Another browser or device will not see it, and private browsing, storage restrictions, or clearing site data can remove it.
Check that you are using the same browser profile and device. If site storage was cleared or blocked, OpenGPT cannot restore the local list.
Use real tasks and keep the decision project on this device.
Compare two specific versions with a template or your own editable tasks.
Choose two different catalog model IDs that support the same-task workflow. Open both provider services, check versions and settings, remove sensitive data, and remember that provider policies apply to prompts submitted there.
The overall preference determines each task result. Close results or fewer than three completed tasks remain inconclusive.
Hiding labels only reduces visible-label bias. It is not double-blind, does not independently run models, and does not verify answer origin or establish a universal ranking.
A source record without a provider-documented model ID cannot enter the catalog comparison. The same model may also already be selected, or the four-model list may be full.
Reuse the same real prompts in a later model evaluation without rebuilding the task set.
Use saved task sets to make future comparisons more consistent and to re-evaluate exact model versions after a release or public-data change.
Finish the task titles and prompts, choose a clear task-set name, and use the same browser and device. A task set can contain sensitive work text, so remove private data before saving.
The task-set name, template, language, task titles, and prompts are stored locally for reuse. Re-evaluation starts a clean run while preserving the intended test inputs.
Saved task sets exist only on this device and are not synced or backed up. Answers, model choices, reviews, reports, and API keys are not saved in a task set. Clearing site data can remove all saved sets.
OpenGPT keeps the same tasks and exact model versions, then clears prior answers, scores, judgments, reveal state, and report so the next run is independent of the old result.
Check supported JSON reports for structure, internal consistency, secret patterns, and a local fingerprint.
Use the verifier to inspect an OpenGPT Evidence v1 or v2 package, or a V3 report downloaded from Decision Workspace. Every check runs locally in this browser.
Obtain the original JSON file. For a comparison report, download it from Decision Workspace; complete every case first if you need a completed result rather than a draft.
For Evidence v1/v2, Passed means the package follows supported structure and consistency rules, while Demo marks an example. For a comparison report, Passed means a completed report is internally consistent; Draft means its structure is valid but the comparison is unfinished. Failed means at least one required check needs review.
A pass does not prove that the named model produced the pasted answers, that provider claims are true, or that the result generalizes beyond the recorded tests. Results from unrelated files are not automatically comparable.
No. It confirms only that the supported report is structurally valid and internally consistent. You must still judge the model identity, test quality, outputs, and decision relevance.
Learn what makes a model decision inspectable and reproducible.
Use this page before designing, publishing, or relying on an evaluation. It separates evidence quality from model popularity.
Start with a concrete decision question, named model versions, representative tasks, and a plan for preserving raw outputs.
Evidence v1/v2 mainly checks format and declarations. Decision Workspace V3 supports local recomputation (L3); L4 independent reproduction remains unsupported.
A rigorous process can reduce bias but cannot guarantee that a sample represents every user, language, future version, or production condition.
Match rigor to consequence. A personal trial may need L1; a public or high-stakes claim should seek stronger independent evidence.
Use these definitions when a label is unfamiliar.
Review which requirement excluded each candidate, then relax only a constraint that is genuinely flexible.
Review requirementsThe publisher may use tied ranks or multiple configurations. OpenGPT preserves the original position instead of renumbering the visible rows.
Review the rankingThe source did not publish a comparable value. OpenGPT leaves it unreported instead of treating it as zero; compare only values from the same benchmark scope.
Review Agent evidenceIt is a complete source record rather than an unresolved status. Open the original source; use the model directory when you need provider-documented identity and official access.
Review source identityUse the same browser profile and device. Private mode, blocked storage, or clearing site data can remove the local list.
Open My ModelsData comparison accepts catalog model IDs. The same model may already be selected, or the list may have reached its four-model limit. Choose another catalog version or remove a candidate first.
Open the model directoryReturn to every marked case and provide both complete outputs plus a verdict.
Continue comparisonReturn to Compare, review the newer local draft warning, then keep or reset the draft intentionally.
Review local draftUse the failed checklist item to correct the source record. Do not insert secrets or rewrite outcomes.
Open verifierReport a broken link, incorrect model detail, or issue you could not resolve.