LLM Benchmarks

Source-backed measurements

Compare 100 AI models across overall capability, knowledge, coding and agents—with sources, test dates and transparent comparison rules.

Benchmarks
Models
Filters
Models
Cell view
Data quality
Compare models
Available models: 100
100 of 100 models · 16 benchmarks
Read the table
Find a model · Click a header to sort · Click a score for its source Line: green newest · blue current · ochre after 3 months≡ Shared rank: the lead is smaller than a single solved task#3/13 Window place: ranked only inside its own task windowDate Date label = date basisS·A+ Source authority: A+ highest · A strong · B solid · C limitedT·Max Highest measured thinking16 16 benchmarks: four per area. Choose a category to see all its tests.Ranks and shades use all rankable models in each benchmark, including hidden models. Five bands are used from ten results; smaller cohorts remain neutral. These ranks are derived, not source measurements.ORG Benchmark organizer · PROV Model provider · 3P Official cross-provider comparison credible sources disagree for the same disclosed protocol directly comparable under the same protocol, with possible benchmark contamination different protocolMeasured quality targetsComparable cells: 90% · target ≥85%Fresh essential cells: 86% · target ≥90%Essential columns with ≥60 models: 6/16
Share & export

For each test setup, we show the highest disclosed thinking level. It may not produce the highest score. Levels are not equal compute budgets across providers.

Latest verified refresh Models +0/−0 · Measurements +52 · updated 0 · −52 · 0 visible table cells changedOpen machine-readable change feed
Cell view
Top 10%Top 25%MiddleBottom quarterBottomRanks and colors refer to all models per benchmark. Model filters only hide rows.Highest measured thinking
Current, source-attributed results for 100 language models. Columns are sortable. Comparable values are ranked; SWE-Pro can also rank clearly labelled official system reports, with differing protocols marked ≈.
OverviewKnowledge and reasoningCodingAgents and tools
⁨ECI⁩ rank ↓ · ModelECI, Overview, Epoch Capabilities Index: 82/100 measured models; 82 comparable. Canonical protocol: Epoch ECI snapshot 2026-10-01 · official-composite-fit · Epoch Capabilities Index. Best comparable first.InfoAA Index, Overview, Artificial Analysis Intelligence Index: 58/100 measured models; 58 comparable, 1 sharing a place. Canonical protocol: Artificial Analysis Intelligence Index v4.3 · intelligence-index-v4-3 · Artificial Analysis · Intelligence Index evaluations. Sort.InfoArena, Overview, Arena Text: 80/100 measured models; 80 comparable, 23 sharing a place. Canonical protocol: 2026-09-30 · overall-style-control · Arena Text. Sort.InfoLiveBench, Overview, LiveBench: 55/100 measured models; 55 comparable. Canonical protocol: 2026-06-25 · overall-seven-category · LiveBench. Sort.InfoHLE Full, Knowledge and reasoning, Humanity's Last Exam · Full multimodal: 54/100 measured models; 16 comparable, 38 with a different protocol, 1 sharing a place. Canonical protocol: HLE finalized 2,500 · Scale protocol 2025-04-03 · full-multimodal-2500-no-tools · Scale AI Labs · Humanity's Last Exam Full. Sort.InfoGPQA, Knowledge and reasoning, GPQA Diamond: 89/100 measured models; 56 comparable, 33 with a different protocol, 27 sharing a place. Canonical protocol: GPQA Diamond v1 · Vals dual-prompt · diamond-198-zero-and-five-shot-cot · Vals AI GPQA Diamond. Sort.InfoMMLU-Pro, Knowledge and reasoning, MMLU-Pro: 66/100 measured models; 57 comparable, 9 with a different protocol, 14 sharing a place. Canonical protocol: Vals MMLU-Pro v1 · vals-mmlu-pro-overall · Vals AI · MMLU-Pro. Sort.InfoFrontierMath T1–3, Knowledge and reasoning, FrontierMath T1–3: 64/100 measured models; 64 comparable, 16 sharing a place. Canonical protocol: 2.0.0 · private-285-of-295-problem-tiers-1-3-set · Epoch AI · FrontierMath Tiers 1–3 v2. Sort.InfoSWE-bench, Coding, SWE-bench Verified: 60/100 measured models; 60 comparable, 18 sharing a place. Canonical protocol: SWE-bench Verified v1 · verified-500-vals-mini-swe-agent · Vals AI SWE-bench Verified. Shows only Vals AI's uniform 500-task mini-swe-agent/Bash evaluation. This is the default directly comparable view. Sort.InfoSWE-rebench, Coding, SWE-rebench v2: 38/100 measured models; 13 comparable, 8 directly comparable with a contamination warning, 25 with a different protocol. Canonical protocol: rolling-v2 · 2026-05-15/2026-07-01 · SWE-rebench. Shows the latest shared SWE-rebench task window. Contamination warnings remain visible and do not change the task-window match. 2026-05-15/2026-07-01 Sort.InfoLiveCodeBench, Coding, LiveCodeBench: 63/100 measured models; 55 comparable, 8 with a different protocol, 8 sharing a place. Canonical protocol: Vals LiveCodeBench v1 · vals-livecodebench-overall · Vals AI · LiveCodeBench. Sort.InfoSciCode, Coding, SciCode: 79/100 measured models; 70 comparable, 9 with a different protocol, 8 sharing a place. Canonical protocol: SciCode current leaderboard · scientific-research-coding-subproblems · Epoch AI mirror · SciCode. Sort.InfoTerminal 2.1, Agents and tools, Terminal-Bench 2.1: 67/100 measured models; 58 comparable, 9 with a different protocol, 12 sharing a place. Canonical protocol: Vals Terminal-Bench 2.1 v2.1 · vals-terminal-bench-2-1-overall · Vals AI · Terminal-Bench 2.1. Sort.InfoMCP-Atlas, Agents and tools, MCP-Atlas · Scale: 24/100 measured models; 23 comparable, 1 with a different protocol, 2 sharing a place. Canonical protocol: MCP-Atlas April 2026 rescored protocol · all-1000-public-500-private-500 · Scale AI Labs · MCP-Atlas. Sort.InfoAPEX, Agents and tools, APEX-Agents: 53/100 measured models; 53 comparable, 3 sharing a place. Canonical protocol: APEX-Agents current leaderboard · 480-professional-tasks-33-worlds · Epoch AI mirror · APEX-Agents. Sort.InfoGDPval-AA v2.1, Agents and tools, GDPval-AA v2.1: 76/100 measured models; 76 comparable, 2 sharing a place. Canonical protocol: GDPval-AA v2.1 (Crowd-BT Elo, DeepSeek V4.1 Flash (max) at 1600) · 220-tasks-elo-v2-1 · Artificial Analysis · Intelligence Index evaluations. Sort.Info
Rank 1Claude Opus 5.5Anthropic10/16 measured · 9 comparable
Rank 2GPT 6 astraOpenAI11/16 measured · 10 comparable
Rank 3Claude Sonnet 5.5Anthropic9/16 measured · 8 comparable
Rank 4Claude Fable 5.1Anthropic13/16 measured · 12 comparable
Rank 5Claude Opus 5Anthropic16/16 measured · 15 comparable
Rank 6GPT 5.5 ProOpenAI3/16 measured · 2 comparable
Rank 7Claude Fable 5Anthropic15/16 measured · 15 comparable
Rank 8GPT 5.6 SolOpenAI15/16 measured · 15 comparable
Rank 9GPT 5.6 TerraOpenAI13/16 measured · 13 comparable
Rank 10GPT 5.5OpenAI15/16 measured · 13 comparable
Rank 11GPT 5.4 ProOpenAI4/16 measured · 3 comparable
Rank 12Claude Opus 4.8Anthropic15/16 measured · 13 comparable
Rank 13Kimi K3Moonshot AI15/16 measured · 14 comparable
Rank 14Gemini 3.7 FlashGoogle13/16 measured · 13 comparable
Rank 15Gemini 3.8 FlashGoogle14/16 measured · 14 comparable
Rank 16GPT 5.4OpenAI14/16 measured · 13 comparable
Rank 17Muse Spark 1 3Meta10/16 measured · 9 comparable
Rank 18GPT 5.3 CodexOpenAI7/16 measured · 5 comparable
Rank 19Grok 4.6xAI13/16 measured · 13 comparable
Rank 20Qwen 3.8 MaxAlibaba12/16 measured · 12 comparable
Rank 21GPT 5.6 LunaOpenAI12/16 measured · 12 comparable
Rank 22Claude Opus 4.7Anthropic15/16 measured · 14 comparable
Rank 23Claude Sonnet 5Anthropic15/16 measured · 14 comparable
Rank 24GLM 5.3Z.ai14/16 measured · 13 comparable
Rank 25GPT 5.2 ProOpenAI2/16 measured · 2 comparable
Rank 26DeepSeek V4 Pro (2026-08-13)DeepSeek14/16 measured · 13 comparable
Rank 27Claude Opus 4.6Anthropic12/16 measured · 11 comparable
Rank 28Qwen 3.8 Max (0902)Alibaba6/16 measured · 5 comparable
Rank 29DeepSeek V4.1 FlashDeepSeek8/16 measured · 7 comparable
Rank 30Muse Spark 1.2Meta11/16 measured · 10 comparable
Rank 31Gemini 3.1 ProGoogle16/16 measured · 15 comparable
Rank 32DeepSeek V4 Flash (0731)DeepSeek12/16 measured · 11 comparable
Rank 33Gemini 3.5 FlashGoogle16/16 measured · 14 comparable
Rank 34Gemini 3.6 FlashGoogle13/16 measured · 13 comparable
Rank 35Muse Spark 1.1Meta13/16 measured · 12 comparable
Rank 36Grok 4.5xAI14/16 measured · 14 comparable
Rank 37Qwen 3.7 MaxAlibaba11/16 measured · 11 comparable
Rank 38GPT 5.2OpenAI11/16 measured · 10 comparable
Rank 39Gemini 3 ProGoogle10/16 measured · 9 comparable
Rank 40Claude Sonnet 4.6Anthropic15/16 measured · 13 comparable
Rank 41Muse SparkMeta9/16 measured · 9 comparable
Rank 42Grok 4.20xAI8/16 measured · 8 comparable
Rank 43GLM 5.3 FlashZ.ai14/16 measured · 13 comparable
Rank 44Gemini 3 FlashGoogle11/16 measured · 9 comparable
Rank 45GLM 5.2Z.ai744B · 40B active16/16 measured · 15 comparable
Rank 46Kimi K2.6Moonshot AI1T · 32B active14/16 measured · 12 comparable
Rank 47GPT 5 ProOpenAI3/16 measured · 3 comparable
Rank 48Inkling SmallThinking Machines Lab276B · 12B active13/16 measured · 12 comparable
Rank 49Claude Opus 4.5Anthropic11/16 measured · 10 comparable
Rank 50Kimi K2.7 CodeMoonshot AI1T · 32B active11/16 measured · 10 comparable
Rank 51GPT 5OpenAI11/16 measured · 10 comparable
Rank 52GLM 5.1Z.ai744B · 40B active14/16 measured · 12 comparable
Rank 53GPT 5.1OpenAI10/16 measured · 10 comparable
Rank 54Qwen 3.8 27BAlibaba27B12/16 measured · 11 comparable
Rank 55Qwen 3.6 Max PreviewAlibaba4/16 measured · 3 comparable
Rank 56Grok 4.3xAI13/16 measured · 12 comparable
Rank 57DeepSeek V4 Pro (2026-04-22)DeepSeek1.6T · 49B active14/16 measured · 13 comparable
Rank 58GPT 5.4 MiniOpenAI12/16 measured · 12 comparable
Rank 59InklingThinking Machines Lab975B · 41B active15/16 measured · 14 comparable
Rank 60Kimi K2.5Moonshot AI1T · 32B active12/16 measured · 11 comparable
Rank 61Qwen 3.6 PlusAlibaba11/16 measured · 11 comparable
Rank 62Qwen 3.7 PlusAlibaba9/16 measured · 7 comparable
Rank 63MiniMax M3MiniMax14/16 measured · 13 comparable
Rank 64Claude Sonnet 4.5Anthropic12/16 measured · 11 comparable
Rank 65Qwen 3.5 397B-A17BAlibaba397B · 17B active12/16 measured · 6 comparable
Rank 66Qwen 3.6 27BAlibaba27B13/16 measured · 8 comparable
Rank 67DeepSeek V4 Flash (2026-04-22)DeepSeek284B · 13B active9/16 measured · 5 comparable
Rank 68GLM 5Z.ai9/16 measured · 7 comparable
Rank 69GPT 5.4 NanoOpenAI12/16 measured · 12 comparable
Rank 70Gemini 2.5 ProGoogle9/16 measured · 7 comparable
Rank 71Gemini 3.5 Flash LiteGoogle12/16 measured · 12 comparable
Rank 72Gemini 3.1 Flash LiteGoogle13/16 measured · 13 comparable
Rank 73Claude Opus 4.1Anthropic7/16 measured · 6 comparable
Rank 74Qwen 3.6 35B-A3BAlibaba35B · 3B active11/16 measured · 6 comparable
Rank 75Gemma 4 31B ITGoogle30.7B6/16 measured · 2 comparable
Rank 76Qwen 3.5 35B-A3BAlibaba35B · 3B active11/16 measured · 6 comparable
Rank 77GPT 5.5 InstantOpenAI5/16 measured · 4 comparable
Rank 78mistral Medium 3.5Mistral AI9/16 measured · 8 comparable
Rank 79Qwen 3.5 9BAlibaba9B8/16 measured · 4 comparable
Rank 80Qwen 3 32BAlibaba32.8B5/16 measured · 4 comparable
Rank 81Qwen 3 14BAlibaba14.8B4/16 measured · 3 comparable
Rank 82GPT 4.5OpenAI2/16 measured · 2 comparable
18 models without a directly comparable ECI score
No rankGemini 4 argonGoogle4/16 measured · 3 comparable
No rankGPT 6 SolOpenAI9/16 measured · 8 comparable
No rankGrok 4.7xAI6/16 measured · 6 comparable
No rankGPT 6 LunaOpenAI9/16 measured · 8 comparable
No rankMiMo V2.6 ProXiaomi5/16 measured · 4 comparable
No rankGLM 5.3 MaxZ.ai1/16 measured · 1 comparable
No rankGPT 5.2 Chat Latest (2026-02-10)OpenAI1/16 measured · 1 comparable
No rankMiMo V2.5 ProXiaomi11/16 measured · 10 comparable
No rankHY3Tencent7/16 measured · 3 comparable
No rankGrok 4.1xAI2/16 measured · 2 comparable
No rankQwen 3.5 27BAlibaba27B6/16 measured · 1 comparable
No rankQwen 3.8 Flash NextAlibaba125B · 6B active7/16 measured · 3 comparable
No rankGemma 4 12B ITGoogle12B4/16 measured · 0 comparable
No rankGemma 4 26B-A4B ITGoogle25.2B · 3.8B active4/16 measured · 0 comparable
No rankQwen 3.8 2.4T-A95BAlibaba2.4T · 95B active6/16 measured · 3 comparable
No rankMiMo V2.6 FlashXiaomi5/16 measured · 4 comparable
No rankGPT 6.1 SolOpenAI7/16 measured · 6 comparable
No rankGemma 4 31BGoogle7/16 measured · 3 comparable