公開榜單來源
LiveBench AI 模型榜單資料
OpenGPT 在原始來源範圍內保留22筆 LiveBench 公開記錄,不重新測試,也不與其他評測合成總分。
- 公開記錄
- 22
- 榜單維度
- 客觀工作 · 單位得分成本 · 開放權重
- 來源資料日期
- 資料取得日期
資料範圍
LiveBench-2026-06-25. Twenty-three objective tasks across seven categories; questions refresh every six months.
OpenGPT 如何處理這些資料
OpenGPT 在同一公開範圍內保存發布方的名稱、名次、設定、日期與數值;不把缺失值改成零,也不產生跨來源總分。
可比範圍
- 名次和數值只能在指定來源版本與榜單範圍內比較。
- 只有在模型身分對應依據經過核驗後,來源設定才會連結到官方模型 ID。
- 公開榜單用於篩選候選;正式採用前仍需用真實工作測試準確版本。
LiveBench-2026-06-25
展示22筆公開記錄中的14筆代表記錄。進入單筆記錄可核對來源身分與對應依據。
- Claude Fable 5Anthropic · 客觀工作#183.0Max effort
- GPT-5.6 SolOpenAI · 客觀工作#281.1Max effort
- GPT-5.5OpenAI · 客觀工作#380.2xHigh effort
- Claude Opus 5Anthropic · 客觀工作#480.1Max effort
- Kimi K3Moonshot AI · 客觀工作#579.2Published benchmark configuration
- Qwen 3.8Alibaba · 客觀工作#678.5Max effort
- DeepSeek V4 FlashDeepSeek · 單位得分成本#1$0.0161Published benchmark configuration
- Grok Build 0.1xAI · 單位得分成本#2$0.0240Published benchmark configuration
- DeepSeek V4 ProDeepSeek · 單位得分成本#3$0.0498Published benchmark configuration
- DeepSeek V4 Flash 0731DeepSeek · 單位得分成本#4$0.0596Published benchmark configuration
- MiniMax M3MiniMax · 單位得分成本#5$0.0597Published benchmark configuration
- Grok 4.3xAI · 單位得分成本#6$0.0614Published benchmark configuration
- Kimi K3Moonshot AI · 開放權重#179.2Open weights
- DeepSeek V4 Flash 0731DeepSeek · 開放權重#274.2Open weights