Public ranking sources

LiveBench AI model ranking data

OpenGPT preserves 22 published LiveBench records in their original source scope. It does not rerun or blend this ranking with another benchmark.

Published records
22
Ranking views
Objective tasks · Cost per score point · Open-weight models
Source data date
Retrieved

Evidence scope

LiveBench-2026-06-25. Twenty-three objective tasks across seven categories; questions refresh every six months.

How OpenGPT handles this evidence

OpenGPT stores the publisher's labels, positions, configurations, dates, and values within the same published scope. It does not convert missing values to zero or merge this source into a cross-source score.

Comparison boundaries

  • Ranks and values are comparable only inside the named source release and ranking scope.
  • A source-listed configuration is linked to an official model ID only when identity evidence has been reviewed.
  • Published rankings are screening signals; production choice still requires testing the exact version on real work.

LiveBench-2026-06-25

Showing 14 representative rows from 22 published records. Follow a row to inspect its exact source identity and mapping evidence.

  1. Claude Fable 5Anthropic · Objective tasks#183.0Max effort
  2. GPT-5.6 SolOpenAI · Objective tasks#281.1Max effort
  3. GPT-5.5OpenAI · Objective tasks#380.2xHigh effort
  4. Claude Opus 5Anthropic · Objective tasks#480.1Max effort
  5. Kimi K3Moonshot AI · Objective tasks#579.2Published benchmark configuration
  6. Qwen 3.8Alibaba · Objective tasks#678.5Max effort
  7. DeepSeek V4 FlashDeepSeek · Cost per score point#1$0.0161Published benchmark configuration
  8. Grok Build 0.1xAI · Cost per score point#2$0.0240Published benchmark configuration
  9. DeepSeek V4 ProDeepSeek · Cost per score point#3$0.0498Published benchmark configuration
  10. DeepSeek V4 Flash 0731DeepSeek · Cost per score point#4$0.0596Published benchmark configuration
  11. MiniMax M3MiniMax · Cost per score point#5$0.0597Published benchmark configuration
  12. Grok 4.3xAI · Cost per score point#6$0.0614Published benchmark configuration
  13. Kimi K3Moonshot AI · Open-weight models#179.2Open weights
  14. DeepSeek V4 Flash 0731DeepSeek · Open-weight models#274.2Open weights