LMArena
The complete latest style-controlled overall ranking, based on users comparing two model answers side by side.
- Leaderboard publication date
- Jul 21, 2026
- Data retrieved
- Jul 24, 2026
- License
- CC BY 4.0
July 2026 · Research edition
Compare exact model configurations across independent user-preference, intelligence, objective-task, cost, and open-weight-model rankings—then test two on your own work.
OpenGPT organizes and presents publicly released third-party evaluation data; it did not rerun these model tests. Use the rankings to understand relative performance, then validate candidates on your own work.
Ranking methodJuly 2026 · Research edition
Start with the decision you need to make. Each scenario opens the most relevant published ranking signal. The original source and scale remain visible.
Show every publisher record with its source model ID or source name, including aliases and run configurations. Source ranks are preserved; ties can still cause rank numbers to skip.
LMArena’s complete latest style-controlled overall ranking. Every publisher record keeps its source model ID; a provider-documented model ID appears separately when available.
Compared with previous edition Jun 25, 2026. Same source and method. Rank movement is relative; it does not by itself prove a change in model quality.
Showing 40 of 378. User preference score.
Ranking basis
Ranking basis
Scores stay on their original publisher scales. They are not combined into a composite score.
Missing data does not mean zero. It means the publisher did not report that metric for this snapshot.
When a source reports model settings, the position applies to that exact setup and source release—not every task, region, or deployment.
No provider can pay for placement. OpenGPT links directly to the evidence used.
Ranking basis
The complete latest style-controlled overall ranking, based on users comparing two model answers side by side.
Nine evaluations across agentic work, coding, general knowledge, and scientific reasoning.
Twenty-three objective tasks across seven categories; questions refresh every six months.
Method
Review new model and benchmark releases; flag material changes without silently rewriting this edition.
Publish a dated snapshot with exact configurations, source releases, and corrections.
Review categories, source quality, conflicts, and inclusion rules before changing the method.