Public ranking sources

LiveBench AI model ranking data

OpenGPT preserves 30 published LiveBench records in their original source scope. It does not rerun or blend this ranking with another benchmark.

Published records
30
Ranking views
Objective tasks · Cost per score point · Open-weight models
Source data date
Retrieved

Evidence scope

LiveBench-2026-06-25. 46 published model configurations across twenty-three objective tasks and seven categories; questions refresh every six months.

How OpenGPT handles this evidence

OpenGPT stores the publisher's labels, positions, configurations, dates, and values within the same published scope. It does not convert missing values to zero or merge this source into a cross-source score.

Comparison boundaries

  • Ranks and values are comparable only inside the named source release and ranking scope.
  • A source-listed configuration is linked to an official model ID only when identity evidence has been reviewed.
  • Published rankings are screening signals; production choice still requires testing the exact version on real work.

LiveBench-2026-06-25

Showing 18 representative rows from 30 published records. Follow a row to inspect its exact source identity and mapping evidence.

  1. Claude Fable 5 Max EffortAnthropic · Objective tasks#182.97Published benchmark configuration
  2. GPT-5.6 Sol Max EffortOpenAI · Objective tasks#281.05Published benchmark configuration
  3. GPT-5.5 Thinking xHigh EffortOpenAI · Objective tasks#380.19Published benchmark configuration
  4. Claude 5 Opus Thinking Max EffortAnthropic · Objective tasks#480.08Published benchmark configuration
  5. Kimi K3Moonshot AI · Objective tasks#579.19Published benchmark configuration
  6. Gemini 3.7 Flash HighGoogle · Objective tasks#678.83Published benchmark configuration
  7. Grok Build 0.1xAI · Cost per score point#1$0.0240Published benchmark configuration
  8. GLM-5.3 FlashZ.AI · Cost per score point#2$0.0305Published benchmark configuration
  9. Qwen 3.8 Flash NextAlibaba · Cost per score point#3$0.0423Published benchmark configuration
  10. DeepSeek V4 Pro 0813DeepSeek · Cost per score point#4$0.0442Published benchmark configuration
  11. DeepSeek V4 Flash Vision ExpDeepSeek · Cost per score point#5$0.0505Published benchmark configuration
  12. DeepSeek V4 Flash 0731DeepSeek · Cost per score point#6$0.0596Published benchmark configuration
  13. Kimi K3Moonshot AI · Open-weight models#179.19Open weights
  14. Qwen 3.8 MaxAlibaba · Open-weight models#278.46Open weights
  15. DeepSeek V4 Pro 0813DeepSeek · Open-weight models#377.44Open weights
  16. Qwen 3.8 Flash NextAlibaba · Open-weight models#476.19Open weights
  17. GLM-5.3Z.AI · Open-weight models#576.14Open weights
  18. Qwen3.8 27BAlibaba · Open-weight models#675.27Open weights