AGI Ranker
A free leaderboard that blends 10 independent AI benchmarks into a single, transparent score so you can compare frontier AI models without wading through a dozen separate charts.
🔗 Visit AGI RankerDescription
Every AI lab publishes its own benchmark numbers, and it's genuinely hard to know which model is actually better once you're staring at ten different scores that don't agree with each other. AGI Ranker's job is to do that comparison work for you: it pulls together results from ten independent, well-known benchmarks — things like GPQA Diamond, SWE-bench, and ARC-AGI-2 — and combines them into one composite score per model, weighted by how much each benchmark says about real capability (agency, reasoning, knowledge, and so on).
What makes it worth trusting rather than just another leaderboard is the methodology: sources are split into four quality tiers, with independent, third-party benchmark results counted at full weight and a lab's own self-reported numbers weighted down or excluded entirely — a direct hedge against labs marking their own homework. It publishes a public corrections log documenting every methodology change, follows a zero-imputed-values policy (no guessing at missing data), and offers interactive tools including an adjustable-weight model explorer and a value-per-dollar API cost comparison. All of it is free, with the underlying data exportable under a CC BY 4.0 license.
💬 Our review
The short version: AGI Ranker is a genuinely useful antidote to benchmark fatigue — instead of trusting whichever number a lab chose to headline, you get a transparent, source-weighted composite that's explicit about what it trusts and why, and it's free.
Against LMSYS Chatbot Arena, which ranks models by crowdsourced human preference in blind chat comparisons, AGI Ranker takes the opposite approach — hard benchmark scores rather than subjective preference — so the two are actually complementary rather than competing. Against the Hugging Face Open LLM Leaderboard, AGI Ranker is narrower in model count (~21 tracked) but deeper in methodology, with its tiered source-quality system and public corrections log giving it more credibility per model covered. The real caveat is churn: methodology has been under continuous refinement since May 2026, which is good for accuracy but makes historical trend-lines shakier if you're trying to track how a model's relative standing moved over time. For a snapshot comparison of frontier models today, though, it's one of the more rigorous free tools available.
💰 Pricing
📊 Global score
🤖 AI-enriched data
Gratuit, sans abonnement. Données exportables sous licence CC BY 4.0.
Pros
Score composite transparent agrégeant 10 benchmarks indépendants avec pondération publiée
Système de tiers de qualité des sources, priorise la vérification indépendante sur l'auto-déclaration des labs
Journal public des corrections, données exportables gratuitement (CC BY 4.0)
Cons
Couverture plus restreinte (~21 modèles) que certains leaderboards grand public
Méthodologie encore fréquemment révisée, rendant les tendances historiques moins stables
Analyse valeur/prix limitée aux modèles avec tarification API publique
