AI model rankings

Submit model

Compare exact model configurations across independent user-preference, intelligence, objective-task, cost, and open-weight-model rankings—then test two on your own work.

Published evaluation data3 external public sourcesEdition data cutoff 2026-07

OpenGPT organizes published third-party evaluation data; it did not rerun these tests. Use the rankings to shortlist models, then validate them on your own work.

Choose a ranking metric

Start with the decision you need to make. Each scenario opens the most relevant published ranking signal. The original source and scale remain visible.

Choose what to show

Show every publisher record with its source model ID or source name, including aliases and run configurations. Source ranks are preserved; ties can still cause rank numbers to skip.

User preference score. LMArena’s complete latest style-controlled overall ranking. Every publisher record keeps its source model ID; a provider-documented model ID appears separately when available.

Compared with previous edition Jun 25, 2026. Same source and method. Rank movement is relative; it does not by itself prove a change in model quality.

Showing 20 of 378. User preference score.

  1. Rank 1
    Model ID
    claude-fable-5
    Votes cast
    14,646
    License type
    Closed / proprietaryProprietary
    User preference score1507Based on anonymous head-to-head comparisons
    Likely score range1,500.91,513.7
  2. Rank 2
    Model ID
    claude-opus-4-6
    Votes cast
    63,191
    License type
    Closed / proprietaryProprietary
    User preference score1505Based on anonymous head-to-head comparisons
    Likely score range1,501.11,508.6
  3. Rank 3
    Model ID
    claude-opus-4-7
    Votes cast
    50,683
    License type
    Closed / proprietaryProprietary
    User preference score1502Based on anonymous head-to-head comparisons
    Likely score range1,497.81,506.2
  4. Rank 4
    Model ID
    claude-opus-4-6
    Votes cast
    67,037
    License type
    Closed / proprietaryProprietary
    User preference score1498Based on anonymous head-to-head comparisons
    Likely score range1,494.11,501.4
  5. Rank 5
    Model ID
    muse-spark-1.1
    Votes cast
    7,927
    License type
    Closed / proprietaryProprietary
    User preference score1495Based on anonymous head-to-head comparisons
    Likely score range1,487.81,502.3
  6. Rank 6
    Model ID
    claude-opus-4-7
    Votes cast
    51,788
    License type
    Closed / proprietaryProprietary
    User preference score1494Based on anonymous head-to-head comparisons
    Likely score range1,489.51,497.9
  7. Rank 7
    Model ID
    muse-spark
    Votes cast
    13,565
    License type
    Closed / proprietaryProprietary
    User preference score1488Based on anonymous head-to-head comparisons
    Likely score range1,481.71,493.4
  8. Rank 8
    Model ID
    gemini-3.1-pro-preview
    Votes cast
    84,631
    License type
    Closed / proprietaryProprietary
    User preference score1486Based on anonymous head-to-head comparisons
    Likely score range1,482.31,489.3
  9. Rank 9
    Source model ID
    gemini-3-pro
    Votes cast
    41,268
    License type
    Closed / proprietaryProprietary
    User preference score1486Based on anonymous head-to-head comparisons
    Likely score range1,481.91,489.6
    Model detailsView source
  10. Rank 10

    kimi-k3

    Moonshot AI
    Model ID
    kimi-k3
    Votes cast
    3,619
    License type
    Closed / proprietaryProprietary
    User preference score1486Based on anonymous head-to-head comparisons
    Likely score range1,475.71,495.6
  11. Rank 11
    Model ID
    gpt-5.6-sol
    Votes cast
    6,221
    License type
    Closed / proprietaryProprietary
    User preference score1485Based on anonymous head-to-head comparisons
    Likely score range1,477.21,493.0
  12. Rank 12
    Model ID
    gemini-3.6-flash
    Votes cast
    4,747
    License type
    Closed / proprietaryProprietary
    User preference score1485Based on anonymous head-to-head comparisons
    Likely score range1,476.21,493.9
  13. Rank 13
    Model ID
    claude-opus-4-8
    Votes cast
    30,901
    License type
    Closed / proprietaryProprietary
    User preference score1484Based on anonymous head-to-head comparisons
    Likely score range1,478.51,488.7
  14. Rank 14
    Source model ID
    gpt-5.5-high
    Votes cast
    45,760
    License type
    Closed / proprietaryProprietary
    User preference score1482Based on anonymous head-to-head comparisons
    Likely score range1,477.31,486.1
    Model detailsView source
  15. Rank 15
    Model ID
    gpt-5.4
    Votes cast
    58,997
    License type
    Closed / proprietaryProprietary
    User preference score1478Based on anonymous head-to-head comparisons
    Likely score range1,473.81,481.7
  16. Rank 16
    Source model ID
    gpt-5.5
    Votes cast
    47,180
    License type
    Closed / proprietaryProprietary
    User preference score1476Based on anonymous head-to-head comparisons
    Likely score range1,472.01,480.7
    Model detailsView source
  17. Rank 17
    Model ID
    gemini-3.5-flash
    Votes cast
    10,092
    License type
    Closed / proprietaryProprietary
    User preference score1476Based on anonymous head-to-head comparisons
    Likely score range1,469.71,482.8
  18. Rank 18
    Source model ID
    gpt-5.2-chat-latest-20260210
    Votes cast
    34,420
    License type
    Closed / proprietaryProprietary
    User preference score1476Based on anonymous head-to-head comparisons
    Likely score range1,471.61,479.8
    Model detailsView source
  19. Rank 19
    Source model ID
    qwen3.7-max-preview
    Votes cast
    3,714
    License type
    Closed / proprietaryProprietary
    User preference score1475Based on anonymous head-to-head comparisons
    Likely score range1,465.11,485.2
    Model detailsView source
  20. Rank 20
    Source model ID
    grok-4.20-beta1
    Votes cast
    26,822
    License type
    Closed / proprietaryProprietary
    User preference score1474Based on anonymous head-to-head comparisons
    Likely score range1,469.71,479.0
    Model detailsView source
Showing 20 of 378
Independent sourceLMArena378 source-listed model entries · Data retrieved: Jul 24, 2026
LMArena

Monitoring statusLive status unavailableValidated automatic publication

Source data date
OpenGPT retrieved
Ranking last changed

LMArena: validated automatic publication · Artificial Analysis and LiveBench: manual review

Ranking methodEach view preserves the source’s metric, release, and methodology. Follow the links to audit the underlying work.

Ranking basis

Full published ranking and official model IDs

The main ranking preserves every publisher record and shows its source model ID or source name. Every row can be compared: exact, configuration-safe mappings use official versions; all other rows stay within one source ranking. Missing data never becomes a zero score.
79 official model IDs
79
43 with public ranking data
43
36 awaiting comparable data
36
378 source records retained for audit
378
Model directory

Ranking basis

How to read this research edition

  • Scores stay on their original publisher scales. They are not combined into a composite score.

  • Missing data does not mean zero. It means the publisher did not report that metric for this snapshot.

  • When a source reports model settings, the position applies to that exact setup and source release—not every task, region, or deployment.

  • No provider can pay for placement. OpenGPT links directly to the evidence used.

Ranking basis

Data and sources

Each view preserves the source’s metric, release, and methodology. Follow the links to audit the underlying work.
AR

LMArena

The complete latest style-controlled overall ranking, based on users comparing two model answers side by side.

Leaderboard publication date
Jul 21, 2026
Data retrieved
Jul 24, 2026
License
CC BY 4.0

Original rankingScoring method

AA

Artificial Analysis

Nine evaluations across agentic work, coding, general knowledge, and scientific reasoning.

Source snapshot
Intelligence Index v4.1 · continuously updated
Data retrieved
Jul 24, 2026

Original rankingScoring method

LB

LiveBench

Twenty-three objective tasks across seven categories; questions refresh every six months.

Source snapshot
LiveBench-2026-06-25
Data retrieved
Jul 24, 2026

Original rankingScoring method

Method version

A predictable publishing rhythm

OpenGPT checks model sources every six hours. Validated LMArena updates can publish automatically; Artificial Analysis and LiveBench require manual review. Failed validation keeps the previous ranking active.
  1. 01

    6-hour source checks

    LMArena updates publish automatically only after schema, completeness, date, and anomaly checks. Artificial Analysis and LiveBench stay under manual review.

  2. 02

    Weekly change summary

    Summarize meaningful ranking, model, and source changes in one reviewable update.

  3. 03

    Monthly history snapshot

    Freeze a dated snapshot with exact configurations, source releases, and corrections.