Independent model decisions

AI model rankings

Shortlist exact model versions from public evidence, compare them on one basis, then verify the choice on your own task.

Objective-task scoreLiveBench · Showing 3 of 10
  1. Rank 1Claude Fable 5.1 Max EffortAnthropic83.41
  2. Rank 2Claude Fable 5 Max EffortAnthropic82.97
  3. Rank 3GPT-6 Astra Max EffortOpenAI82.16
Published evaluation data2 external public sourcesLiveBench · Edition data cutoff

OpenGPT organizes published third-party evaluation data; it did not rerun these tests. Use the rankings to shortlist models, then validate them on your own work.

Data edition
2026-09Published ·
Automated decision scope
20 in automated verification / 79 catalog modelsData cutoff · 2026-08-30
Current ranking snapshot · LiveBench
2026-09-07 · 56 source records
Latest comparable change edition · LiveBench
Unknown
Lifecycle coverage
36 lifecycle states documented / 79 · 20 in automation scopeData cutoff · 2026-08-30

Choose a ranking metric

Start with the decision you need to make. Each scenario opens the most relevant published ranking signal. The original source and scale remain visible.

Choose what to show

Show every publisher record with its source model ID or source name, including aliases and run configurations. Source ranks are preserved; ties can still cause rank numbers to skip.

Objective-task score. LiveBench overall score across 23 objectively graded tasks in seven categories.

Showing 10 of 10. Objective-task score.

  1. Rank 1
    Model settings
    Published benchmark configuration
    Source model ID
    claude-fable-5-1-max-effort
    Objective-task score
    83.41
    View source
  2. Rank 2
    Model settings
    Published benchmark configuration
    Source model ID
    claude-fable-5-max-effort
    Objective-task score
    82.97
    View source
  3. Rank 3
    Model settings
    Published benchmark configuration
    Source model ID
    gpt-6-astra-max
    Objective-task score
    82.16
    View source
  4. Rank 4
    Model settings
    Published benchmark configuration
    Source model ID
    muse-spark-1.3-xhigh
    Objective-task score
    81.59
    View source
  5. Rank 5
    Model settings
    Published benchmark configuration
    Source model ID
    gpt-5.6-sol-max
    Objective-task score
    81.05
    View source
  6. Rank 6
    Model settings
    Published benchmark configuration
    Source model ID
    gpt-5.5-xhigh
    Objective-task score
    80.19
    View source
  7. Rank 7
    Model settings
    Published benchmark configuration
    Source model ID
    claude-opus-5-max-effort
    Objective-task score
    80.08
    View source
  8. Rank 8

    Kimi K3

    Moonshot AI
    Model settings
    Published benchmark configuration
    Source model ID
    kimi-k3
    Objective-task score
    79.19
    View source
  9. Rank 9
    Model settings
    Published benchmark configuration
    Source model ID
    gemini-3.7-flash-high
    Objective-task score
    78.83
    View source
  10. Rank 10
    Model settings
    Published benchmark configuration
    Source model ID
    qwen3.8-max
    Objective-task score
    78.46
    View source
Independent sourceLiveBench10 source-listed model entries · Data retrieved: Sep 9, 2026
LiveBench

Monitoring statusLive status unavailableValidated automatic publication

Source data date
OpenGPT retrieved
Ranking last changed

LMArena + LiveBench: validated automatic publication · Paused sources: historical only

Ranking methodEach view preserves the source’s metric, release, and methodology. Follow the links to audit the underlying work.

Ranking basis

Full published ranking and official model IDs

The main ranking preserves every publisher record and shows its source model ID or source name. Every row can be compared: exact, configuration-safe mappings use official versions; all other rows stay within one source ranking. Missing data never becomes a zero score.
79 official model IDs
79
0 with public ranking data
0
79 awaiting comparable data
79
LiveBench · 56 source records retained for audit
56
Model directory

Ranking basis

How to read this research edition

  • Scores stay on their original publisher scales. They are not combined into a composite score.

  • Missing data does not mean zero. It means the publisher did not report that metric for this snapshot.

  • When a source reports model settings, the position applies to that exact setup and source release—not every task, region, or deployment.

  • No provider can pay for placement. OpenGPT links directly to the evidence used.

Ranking basis

Data and sources

Each view preserves the source’s metric, release, and methodology. Follow the links to audit the underlying work.
AR

LMArena

The complete latest style-controlled overall ranking, based on users comparing two model answers side by side.

Leaderboard publication date
Sep 2, 2026
Data retrieved
Sep 9, 2026
License
CC BY 4.0

Original rankingScoring method

LB

LiveBench

Twenty-three objective tasks across seven categories; questions refresh every six months.

Source snapshot
LiveBench-2026-06-25
Data retrieved
Sep 9, 2026
License
Apache-2.0 project license; exact new-livebench publication repository does not repeat a LICENSE file

Original rankingScoring method

Method version

A predictable publishing rhythm

OpenGPT checks active model sources every six hours. LMArena and LiveBench publish only after their versioned validation policies pass. Paused sources remain historical and do not enter current rankings.
  1. 01

    6-hour source checks

    LMArena and LiveBench updates publish automatically only after schema, completeness, date, integrity, and anomaly checks. Paused sources stay historical.

  2. 02

    Weekly change summary

    Summarize meaningful ranking, model, and source changes in one reviewable update.

  3. 03

    Monthly history snapshot

    Freeze a dated snapshot with exact configurations, source releases, and corrections.