Loading leaderboard…
AI Model Leaderboard
Performance metrics across various AI models and benchmarks
Frequently asked questions
What do the benchmarks mean?
- Intelligence Index
- Aggregated Artificial Analysis Intelligence Index combining multiple performance metrics into one overall score.
- Coding Index
- Composite Artificial Analysis index focused on programming and software-engineering ability.
- Agentic Index
- Composite Artificial Analysis index focused on completing multi-step tasks with tools and environments.
- HLE
- Humanity’s Last Exam measures frontier academic knowledge across mathematics, sciences, and humanities.
- GPQA Diamond
- Graduate-level questions in biology, physics, and chemistry that test scientific reasoning.
- SciCode
- Scientist-designed coding problems that test the implementation of research algorithms.
- Terminal-Bench 2.1
- Tests agents on practical tasks completed through a terminal environment.
- CritPt
- Research-level physics problems that require code generation and quantitative reasoning.
Why are some models missing?
Model IDs cannot always be matched to Artificial Analysis entries. Provider names may vary or be uncatalogued, so some of your models might not appear in the leaderboard.
Why do input and output capabilities sometimes differ?
Admins can manually configure which capabilities a model provides for input or output in this system, so configured capabilities may differ from a model’s default feature set.