See how a recommendation becomes a decision.
OpenGPT is an AI Decision Intelligence Platform. It keeps editorial requirement matching, factual source records, real-task evaluation, and published benchmark evidence separate so every result can show its conditions and limits.
The method in one minute
- A rank is relative to one named source and scope—not a universal score.
- Exact model versions and complete Agent configurations are the unit of comparison.
- Unknown or incompatible data stays unknown instead of becoming zero or one combined ranking.
- Local report checks prove structure and internal consistency, not output origin or provider authenticity.
Editorial fit scores · editorial-fit-v1
How 0–10 editorial fit scores are assigned
These eight catalog signals support category-by-category comparison. They are a documented OpenGPT editorial judgment layer, separate from rules-v2 recommendation points and public benchmark evidence.
Editorial product-fit score: 10 indicates stronger fit for the named dimension; dimensions are displayed equally and never combined into an overall winner.
Intelligence
0–10Editorial fit for broad, complex general work using documented capabilities, catalog role, and known product constraints—not an IQ or benchmark score.
Coding
0–10Editorial fit for software planning, implementation, debugging, review, and repository workflows; it is not a measured pass rate.
Reasoning
0–10Editorial fit for multi-step analysis, constraint handling, and structured problem solving—not a reproduced reasoning benchmark.
Writing
0–10Editorial fit for drafting, editing, long-form writing, and documentation workflows—not a measured style or factuality result.
Speed
0–10Editorial fit for responsive iteration based on product positioning and delivery mode—not a measured latency claim.
Cost efficiency
0–10Editorial fit for access-price and operating-cost control using listed tiers and deployment mode; it is not a total-cost calculation or price guarantee.
Privacy control
0–10Editorial fit for deployment and data control using local, hybrid, or cloud classification; it is not a security, compliance, or telemetry audit.
Ecosystem
0–10Editorial fit for available tools, integrations, platforms, and deployment options—not a market-share or integration-quality measurement.
Weights and aggregation
Each dimension has equal display prominence. There is no arithmetic weight, average, total score, or overall winner; a user decides which dimensions matter for the stated conditions.
Human editorial boundary
Editors assign integer 0–10 fit signals from the rubric, provider-confirmed product facts, and catalog classifications. A score is not an exact-version measurement, benchmark result, probability, or provider guarantee.
Sources
Official provider pages support identity, access, published capabilities, deployment options, and listed pricing. OpenGPT best-for tags, strengths, weaknesses, classifications, and all 0–10 values are editorial inputs.
Catalog review
Catalog last reviewed:
Confidence: Medium editorial coverage
Confidence describes completeness of the catalog coverage and explanation—not the probability of strong performance. Provider facts can change and the fit signal has not been independently measured.
Applicability
Use the scores to compare catalog entries by the named dimension under the current product and deployment assumptions. Do not compare score gaps as measured performance differences.
Before adoption, verify current provider facts and test the exact model version, service tier, tools, region, and representative work.
Published method · rules-v2
How OpenGPT recommendations are produced
The Builder turns a local Profile into a deterministic shortlist. The points below rank editorial product fit after hard budget and privacy gates; they do not measure model quality.
Scoring dimensions and additive rule points
Points are conditional weights, not percentages. A hard gate removes an ineligible item before scoring.
Task and goal fit
Model · ToolModels receive +12 for each selected task match and +10 for an exact goal match. Tools receive +10 for each selected use-case match.
User-declared conditions · OpenGPT editorial catalogPrivacy and deployment
Model · ToolHigh privacy excludes non-local models and tools not classified high-privacy. Eligible local models receive +30 and high-privacy tools +12. For medium privacy, hybrid/local/cloud models receive +8/+5/+2.
User-declared conditions · OpenGPT editorial catalogBudget eligibility
Model · ToolItems outside the selected listed-price band are excluded. Eligible models receive +12 and tools +8; a free-profile model with a listed free entry receives another +16.
User-declared conditions · Provider-published facts · OpenGPT editorial catalogProfessional role fit
Model · ToolStudent-model matches receive +10; other supported model-role matches +8. Student-tool matches receive +8; other supported tool-role matches +6.
User-declared conditions · OpenGPT editorial catalogIndustry signals
Model · ToolRecognized industry terms map to task categories. Models receive +5 and tools +4 for each mapped task match. Free text itself is not copied into the workflow.
User-declared conditions · OpenGPT editorial catalogLanguage fit
ModelFor a non-English preference, a multilingual catalog match receives +14 and a non-match −3. For English, a writing or documentation fit receives +3.
User-declared conditions · OpenGPT editorial catalogPlatform fit
Model · ToolModel matches score Web +3, Desktop +10, Terminal +8, API +7, and Mobile +9. Tools receive +7 per platform match or −3 when none match.
User-declared conditions · OpenGPT editorial catalogExperience fit
Model · ToolBeginner model fit is +10 (or −4 for a local-model mismatch), advanced fit +7, and intermediate frontier/hybrid fit +2. Beginner/advanced tool fits are +6/+5.
User-declared conditions · OpenGPT editorial catalogSelection and ties
Model · ToolEligible candidates are sorted by rule points; equal scores use name order. Roles are filled in a fixed sequence, and up to three remaining matched models become alternatives.
Deterministic rules-v2Sources and claim boundaries
Provider-published facts
Official provider pages support identity, access, published capability descriptions, and listed pricing. OpenGPT does not treat those pages as independent performance tests.
OpenGPT editorial catalog
Best-for tags, strengths, privacy/deployment classifications, and product-fit signals are reviewed editorial judgments. They are not benchmark measurements or provider guarantees.
User-declared conditions
Goal, tasks, role, budget, privacy, language, platform, skill, and industry signals come from the user's local profile. They describe applicability, not external evidence.
Deterministic rules-v2
The same normalized profile and catalog snapshot produce the same ordering. The arithmetic is explainable, but determinism does not make editorial inputs factual.
Public benchmark context
Public rankings remain separate evidence. rules-v2 does not add benchmark ranks or scores to recommendation points; use them after shortlisting and preserve each source's scope.
Factual versus editorial
Factual claims must trace to a named provider or benchmark source, an exact identity when available, and a review date.
Fit tags, strengths, classifications, assumptions, confidence, and recommendations are OpenGPT editorial judgments. They must not be presented as measured exact-version results.
Review cadence and confidence
Methodology and active recommendation records are scheduled for review at least every 90 days, plus event-driven review. Each result must use its record review date—not this page's date.
Review is triggered when
- A provider changes a model ID, lifecycle, access path, capability statement, or listed price.
- The recommendation engine, weights, hard gates, or editorial taxonomy changes.
- A correction identifies a material source, identity, pricing, privacy, or applicability error.
Confidence rules
Confidence describes completeness of the fit explanation and its supporting record. It does not express a probability that the model will perform well.
- High
- Exact identity and current factual source are present; material fields are resolved, assumptions are visible, and a qualified alternative is named.
- Medium
- The match is explainable, but one source, condition, or current fact needs confirmation before adoption.
- Low
- Identity, freshness, pricing, privacy, or applicability has a material gap. Treat the item only as a candidate to investigate.
Applicability and conditions
- Applies to the rules-v2 editorial shortlist generated from the current reviewed model and tool catalog.
- The result is conditional on the submitted profile and catalog snapshot; change either and the result can change.
- It is a low-stakes decision aid, not proof of quality, live price, benchmark leadership, safety, legality, or fitness for a regulated use.
Required per-result audit display
An integrated recommendation should expose all seven fields below without requiring the user to infer the evidence boundary.
- Why
- Evidence
- Assumptions
- Last reviewed
- Confidence
- Best alternative
- When not to choose
From a real task to a reviewable conclusion
Define the test
Before comparing, record the task, success criteria, inputs, constraints, exact model versions, settings, and date.
Use the same conditions
Give every candidate the same task and rubric. Hidden labels reduce visible-name bias, but this is not a double-blind or independently run test.
Keep the source data
Save prompts, answers, scores, errors, latency, and hashes. Do not fill missing evidence with estimates.
State the limits
Show uncertainty, sample size, exclusions, conflicts, and what the result does not prove.
Evidence levels
Evidence v1/v2 checks file format and declared consistency. Decision Workspace V3 also recomputes hashes, assignments, mappings, scores, and summaries from the saved inputs. It does not prove independent reproduction at L4.
Format checked
The file follows the declared schema and contains no obvious secret fields.
Record complete
The target is a consistent record of tasks, settings, model declarations, and evaluation date. V1 checks only part of this record.
Recomputable
Decision Workspace V3 includes raw pasted outputs and scoring inputs for strict local recomputation; Evidence v1/v2 does not carry the same fields.
Independently reproduced
Target: a separate runner or signed source reproduces the result under comparable conditions. V1 does not verify this.
How AI agent rankings are compared
Each row represents one published agent configuration in a named benchmark scope. It is not a score for the base model alone.
Compare system configurations
Keep every agent, base-model, scaffold, tool, environment, budget, and date field published by the source together; mark unreported fields as unknown.
Preserve original scales
Rank only inside the same benchmark, version, task rules, and comparable run group. OpenGPT does not create a universal cross-benchmark score.
Name the evidence status
Distinguish records listed by a benchmark publisher, upstream-checked submissions, provider reports, and community runs.
Keep operating evidence
Show cost, latency, steps, and human intervention when the source provides comparable values; leave missing values unknown.
What local validation can establish
- ✓Required fields and supported data types
- ✓Internal consistency of scores, runs, dates, and hashes
- ✓Whether the package declares demo data
- ✓Absence of common API-key and secret fields
What it cannot establish alone
- —That the declared provider or model actually produced the output
- —That a score is free from selection bias or weak test design
- —That one result applies to every task, language, user, or future model version
- —That a model is safe, lawful, or appropriate for a high-stakes decision
Publication principles
No pay-to-rank
Sponsorship may be disclosed, but it cannot change scoring, evidence gates, or placement.
Versions, not brand names
A model series is not a stable test subject. Exact model IDs, settings, region, and date matter.
Real tasks before broad claims
OpenGPT recommends shortlists by requirements, then asks users to test their own representative work.
Uncertainty is a result
Ties, small samples, failed runs, and inconclusive evidence should remain visible.