Source-backed measurements
Guides & methodology
Understand benchmarks, compare measurements fairly and explore the sources behind LLM Benchmarks. Methodology, a benchmark guide and analyses grounded in the data.
Where do the scores come from?
A score becomes useful when you know which model was tested and under what conditions. This guide explains how we select, check and present published results in the comparison table.
Our contribution is assembling and checking the evidence. The linked benchmark operators and model providers perform the underlying evaluations.
1. Which models and sources are included?
The selection combines models from the Epoch Capabilities Index with other demanding evaluations, checked recent provider releases and a dedicated allocation for open models. This includes smaller checkpoints that have limited representation on major leaderboards. It is a sample of the market, not a complete model directory.
We use benchmark-operator results, independent evaluations and documented provider reports. Model identity, source, test version and publication conditions must be traceable. A publicly accessible score is not automatically cleared for republication. Provider reports retain their attribution.
2. When are two scores comparable?
A benchmark name alone is not enough. We account for release, version, task window, metric, unit, tools, agent setup and evaluation procedure. Each benchmark has a defined comparison lane. Results from other setups can add useful evidence, but do not silently receive the same ranking status.
- Rank: position within the benchmark's eligible protocol; not an overall verdict on the model.
- ≈: supplementary evidence under different conditions. Check the measurement details first.
- Conflict: credible sources disagree for the same documented setup. No reliable single rank.
- Empty cell: no matching visible measurement. It means neither zero points nor a failed evaluation.
3. Which thinking level is selected?
Within the same source and test series, we prefer the highest documented fixed level: max, xhigh, high, medium, low, minimal, off. We do not retrospectively pick whichever score is highest. A valid Max run can therefore score below High. Different agents, test versions or task windows are not mixed to obtain a higher effort level.
Adaptive means thinking is enabled, but does not establish a fixed maximum budget. Undisclosed settings remain unknown. Two providers' High labels do not necessarily represent the same computation budget.
4. What do dates and source grades mean?
Evaluation date, publication, dataset version and first observation are different events. If a source does not date an individual run, we may use the first time that exact score was observed. This does not prove a recent evaluation date. Measurement details identify the date basis.
Source authority and evidence quality are evaluated separately. A+, A, B and C are our categories for sources and evidence, not statistical probabilities that a score is correct. A strong source grade does not remove differences in test conditions.
5. What the table cannot establish
Small score gaps are not established performance advantages without an appropriate uncertainty analysis. Percentage points are not relative percentage improvements. A composite index also measures something different from success on a specific task; we do not average the columns into our own overall score.
Quantization and hardware results remain supplementary studies. Different backends, GPUs or task subsets prevent isolating the effect of quantization. Weight-file size, GPU capacity and measured memory usage describe different things. The published hardware tests are not in-house laboratory measurements by LLM Benchmarks.
6. Check evidence and report an error
Open a table cell to inspect its source, setup, date and available thinking variants. Download the data as CSV or JSON to inspect them independently. To report an error, include the model, benchmark, affected score and a link to the original evidence. Changed sources are checked against the publication rules again; confirmed errors are corrected in the dataset.
What does each benchmark measure?
The useful metric depends on your task. A broad capability index, a knowledge test and a coding agent evaluation measure different things. This guide helps you interpret each kind of result.
Start with your task, check the evaluation protocol, then compare the scores.
Broad capability: ECI and composite indices
The Epoch Capabilities Index combines results from different tests into a shared capability scale. It offers an initial view across task domains. Its value is not the percentage of tasks solved. New data can change estimates even for unchanged models.
Our interpretation: use an index to shortlist models, then inspect the individual tests relevant to your work. A specialist assistant can suit your task despite a lower broad index score. Scores from different indices should not be subtracted from each other.
Knowledge and reasoning: GPQA, HLE and mathematics
GPQA contains demanding expert-written questions in biology, physics and chemistry. The table tracks its Diamond subset. Such evaluations provide evidence about problem solving under the stated setup. They do not automatically measure reliability in open-ended research.
HLE text-only and full task sets remain separate. Tool access can change the task. In mathematics, task difficulty matters: a strong score on an older or easier set does not establish performance on newer difficult tasks. For a knowledge application, we would also test your own questions against verifiable sources.
Coding: a single task or an entire agent?
A coding evaluation may test a self-contained programming problem or an agent working in an existing repository. SWE-bench evaluates real software issues. Tools, agent orchestration and the test environment also influence the result.
Our interpretation: when selecting a coding agent, compare documented system setups. A narrower coding test may be more relevant to individual functions. Verified, Pro and changing Rebench task windows are not interchangeable. Success on one set is not a success forecast for your repository.
Agents and tools: the workflow is part of the test
Terminal, tool-use and multi-turn agent evaluations require a model to act within a defined environment. Available APIs, agent setup, time limits and call limits are part of what the score means. For example, the two Banking setups remain separate in our table.
Our interpretation: these results are most useful when the test resembles your workflow. Otherwise important aspects such as permissions, recovery from errors and your own success criteria remain untested. Inspect measurement details and try a small set of representative workflows yourself.
Reading instruction and preference evaluations
IFEval tests objectively verifiable instructions, such as required output constraints. This is useful when your workflow depends on following specific requirements. It does not establish comprehensive factual accuracy or the domain quality of every answer.
Arena leaderboards reflect preferences under their comparison procedure. A preferred answer is not necessarily factually correct. Our interpretation: preference evidence can inform writing-style decisions; verifiable correctness also requires domain checks.
All benchmarks in the comparison table
The directory below links to the original pages and lists our comparison lanes. Their version and task information matters more than similar benchmark names. The grouping aids navigation; it does not weight an overall ranking.
Overview
Epoch Capabilities Index
Composite capability index fitted by Epoch AI across more than 50 benchmarks.
- Comparison source
- Epoch Capabilities Index
- Version / tasks
- Epoch ECI snapshot 2026-10-01 · official-composite-fit
- Metric / unit
- eci · index_points
Artificial Analysis Intelligence Index
Artificial Analysis' composite Intelligence Index; one value per model and reasoning setting. The protocol identifies the methodology version.
- Comparison source
- Artificial Analysis · Intelligence Index evaluations
- Version / tasks
- Artificial Analysis Intelligence Index v4.3 · intelligence-index-v4-3
- Metric / unit
- index_score · points
Arena Text
Human preference ranking for text models.
- Comparison source
- Arena Text
- Version / tasks
- 2026-09-30 · overall-style-control
- Metric / unit
- arena_score · points
LiveBench
Frequently refreshed objective evaluation across broad capabilities.
- Comparison source
- LiveBench
- Version / tasks
- 2026-06-25 · overall-seven-category
- Metric / unit
- category_balanced_average · percent
Knowledge and reasoning
Humanity's Last Exam · Full multimodal
Comparable values use the original 2,500-question full multimodal no-tools protocol; text-only and tool-assisted runs stay separate.
- Comparison source
- Scale AI Labs · Humanity's Last Exam Full
- Version / tasks
- HLE finalized 2,500 · Scale protocol 2025-04-03 · full-multimodal-2500-no-tools
- Metric / unit
- accuracy · percent
Humanity's Last Exam · Text-only
Artificial Analysis' text-only subset: 2,158 of the 2,500 questions, no tools. A different question set, not a weaker run of the full protocol.
- Comparison source
- Artificial Analysis · Intelligence Index evaluations
- Version / tasks
- HLE text-only (Artificial Analysis) · text-only-2158-no-tools
- Metric / unit
- accuracy · percent
SimpleQA Verified
Short-answer factual knowledge and abstention quality.
- Comparison source
- Epoch AI SimpleQA Verified
- Version / tasks
- Epoch independent Inspect runs · verified-1000-epoch-inspect
- Metric / unit
- accuracy · percent
GPQA Diamond
Graduate-level science questions in physics, chemistry and biology.
- Comparison source
- Vals AI GPQA Diamond
- Version / tasks
- GPQA Diamond v1 · Vals dual-prompt · diamond-198-zero-and-five-shot-cot
- Metric / unit
- accuracy · percent
MMLU-Pro
Broad academic multiple-choice knowledge and reasoning across fourteen subjects; Vals AI's independent run over the MMLU-Pro set.
- Comparison source
- Vals AI · MMLU-Pro
- Version / tasks
- Vals MMLU-Pro v1 · vals-mmlu-pro-overall
- Metric / unit
- accuracy · percent
FrontierMath T1–3
Independent Epoch AI runs on 295 expert-written advanced mathematics problems; Tier 4 stays separate.
- Comparison source
- Epoch AI · FrontierMath Tiers 1–3 v2
- Version / tasks
- 2.0.0 · private-285-of-295-problem-tiers-1-3-set
- Metric / unit
- mean_score · percent
ARC-AGI-2
Abstract rule induction over 360 evaluation tasks, with pass@2 reasoning variants kept separate.
- Comparison source
- ARC-AGI-2
- Version / tasks
- v2 · semi-private-120
- Metric / unit
- accuracy · percent
FrontierMath T4
Independent Epoch AI runs on the 41 private research-level problems from the 43-problem FrontierMath Tier 4 v2 set; the two public examples and Tiers 1–3 are not mixed into this column.
- Comparison source
- Epoch AI · FrontierMath Tier 4 v2
- Version / tasks
- 2.0.0 · private-41-of-43-problem-tier-4-set
- Metric / unit
- mean_score · percent
MMMU-Pro
Multimodal expert knowledge and reasoning.
- Comparison source
- Vals AI · MMMU-Pro
- Version / tasks
- Vals MMMU-Pro v1 · vals-mmmu-pro-overall
- Metric / unit
- accuracy · percent
CritPt
Research-level physics reasoning across 71 open-ended challenges graded automatically.
- Comparison source
- Artificial Analysis · Intelligence Index evaluations
- Version / tasks
- CritPt (Artificial Analysis) · 70-challenges-official-grading
- Metric / unit
- accuracy · percent
SimpleBench
Trick-question common-sense reasoning where humans still beat frontier models; average of five runs.
- Comparison source
- Epoch AI mirror · SimpleBench
- Version / tasks
- SimpleBench current leaderboard · official-set-avg-at-5
- Metric / unit
- accuracy_avg_at_5 · percent
Chess Puzzles
Epoch AI's 100 engine-generated chess puzzles with one best move each, run and dated by Epoch.
- Comparison source
- Epoch AI · Chess Puzzles
- Version / tasks
- Epoch AI Chess Puzzles v1 · 100-engine-generated-single-best-move-puzzles
- Metric / unit
- mean_score · percent
ARC-AGI-1
Original ARC-AGI abstraction puzzles on the semi-private set as reported by ARC Prize.
- Comparison source
- Epoch AI mirror · ARC-AGI-1
- Version / tasks
- ARC-AGI-1 current leaderboard · semi-private-evaluation-pass-at-2
- Metric / unit
- pass_at_2_accuracy · percent
IFEval
Instruction following. Provider reports and subset studies retain their scoring and sample-set labels; they do not form a uniform cross-model ranking.
Supplementary results; no shared canonical comparison lane is specified.
Original benchmark page ↗IFBench
Precise instruction following: 300 prompts with 344 verifiable constraints, instruction-level strict accuracy as measured by the Swallow LLM Leaderboard.
- Comparison source
- Swallow LLM Leaderboard
- Version / tasks
- swallow-evaluation-instruct · ifbench-test-300-prompts-344-instructions
- Metric / unit
- inst_level_strict_acc · percent
Coding
SWE-bench Verified
Resolution rate on 500 verified real-world software issues; the primary comparison uses one uniform minimal agent harness.
- Comparison source
- Vals AI SWE-bench Verified
- Version / tasks
- SWE-bench Verified v1 · verified-500-vals-mini-swe-agent
- Metric / unit
- resolved_rate · percent
SWE-rebench v2
Fresh time-windowed software issue resolution.
- Comparison source
- SWE-rebench
- Version / tasks
- rolling-v2 · 2026-05-15/2026-07-01
- Metric / unit
- resolved_rate · percent
SWE-bench-Live · Lite
Continuously refreshed software issue resolution with agent, model and submitted reasoning label kept together.
- Comparison source
- SWE-bench-Live · Lite
- Version / tasks
- SWE-bench-Live Lite · lite-300-python
- Metric / unit
- resolved_rate · percent
DeepSWE
Original long-horizon software engineering tasks with separate harness and reasoning-effort results.
- Comparison source
- DeepSWE official live leaderboard
- Version / tasks
- DeepSWE v1.1 live leaderboard · 113-original-long-horizon-tasks
- Metric / unit
- pass_at_1 · percent
SWE-bench Pro Public · Scale
Scale AI Labs' separate 731-task professional software-engineering benchmark. Results keep the submitted model, agent label, confidence interval and row-specific publication date; they are never mixed with SWE-bench Verified, SWE-rebench or vendor supplemental protocols. A July 2026 OpenAI audit reported substantial task-quality concerns, so SWE-Pro should be read as one imperfect signal rather than an absolute model ranking.
- Comparison source
- Scale AI Labs · SWE-bench Pro Public
- Version / tasks
- Scale public 731-task leaderboard · public-731
- Metric / unit
- resolved_rate · percent
CursorBench
Coding-agent evaluation with explicit reasoning levels plus cost, token and step metadata.
- Comparison source
- Epoch AI mirror · CursorBench
- Version / tasks
- CursorBench current leaderboard · official-current-suite
- Metric / unit
- score · percent
FrontierCode 1.1 · Main
Hard mergeability-focused coding tasks; harness and reasoning effort remain attached to each score.
- Comparison source
- Epoch AI mirror · FrontierCode 1.1
- Version / tasks
- FrontierCode 1.1 · main-100-hardest-tasks
- Metric / unit
- mergeability_rubric_mean_at_5 · percent
IOI
IOI 2024 and 2025 olympiad problems evaluated by Vals AI with olympiad graders.
- Comparison source
- Vals AI · IOI
- Version / tasks
- Vals IOI v2 · vals-ioi-2024-2026-overall
- Metric / unit
- score · percent
LiveCodeBench
Competitive programming problems released after training cutoffs; Vals AI's independent implementation, overall pass@1.
- Comparison source
- Vals AI · LiveCodeBench
- Version / tasks
- Vals LiveCodeBench v1 · vals-livecodebench-overall
- Metric / unit
- pass_at_1 · percent
SciCode
Scientific research coding measured by subproblem accuracy with executable Python tests.
- Comparison source
- Epoch AI mirror · SciCode
- Version / tasks
- SciCode current leaderboard · scientific-research-coding-subproblems
- Metric / unit
- subproblem_accuracy · percent
ALE-Bench
Score-based algorithmic optimisation contests (AtCoder Heuristic); performance rating.
- Comparison source
- Epoch AI mirror · ALE-Bench
- Version / tasks
- ALE-Bench current leaderboard · atcoder-heuristic-contest-set
- Metric / unit
- performance_rating · points
Vibe Code Bench
Building complete web applications from scratch in an agent harness, evaluated by Vals AI.
- Comparison source
- Vals AI · Vibe Code Bench
- Version / tasks
- Vals Vibe Code Bench v1.1 · vals-vibe-code-bench-overall
- Metric / unit
- score · percent
Text Arena · Coding
Human preference rating for generated React and TypeScript web applications.
- Comparison source
- Epoch AI mirror · Text Arena (Coding)
- Version / tasks
- Text Arena Coding current pool · text-arena-coding-bradley-terry
- Metric / unit
- bradley_terry_rating · points
WeirdML
Unusual machine-learning tasks solved end-to-end by writing and running code.
- Comparison source
- Epoch AI mirror · WeirdML
- Version / tasks
- WeirdML v2 current leaderboard · weirdml-v2-task-set
- Metric / unit
- accuracy · percent
Agents and tools
Terminal-Bench 2.0
Terminal-Bench 2.0 leaderboard over 89 hard terminal tasks; agent harness and run date stay attached to every score.
- Comparison source
- Epoch AI mirror · Terminal-Bench 2.0
- Version / tasks
- Terminal-Bench 2.0 · official-89-task-set
- Metric / unit
- accuracy_mean · percent
Terminal-Bench 2.1
Vals AI's independent Terminal-Bench 2.1 run with one uniform harness; vendor reports remain supplemental.
- Comparison source
- Vals AI · Terminal-Bench 2.1
- Version / tasks
- Vals Terminal-Bench 2.1 v2.1 · vals-terminal-bench-2-1-overall
- Metric / unit
- accuracy · percent
Terminal-Bench 3.0
Official agent-plus-model leaderboard over 74 hard terminal tasks; agent, effort and trial count stay attached to every score.
- Comparison source
- Terminal-Bench 3.0
- Version / tasks
- 3.0 · b936386b-4afd-4ad7-aa7f-6b52bcab9853 · official-74-task-leaderboard
- Metric / unit
- accuracy · percent
τ³-Banking · Official harness
Knowledge-intensive multi-turn banking support tasks with retrieval, tools and a simulated user, as published on the official leaderboard.
- Comparison source
- τ³-Banking
- Version / tasks
- 1.0.1 · banking_knowledge
- Metric / unit
- pass_1 · percent
τ³-Banking · Artificial Analysis harness
Artificial Analysis' own run of the same domain: 97 tasks, BM25/grep retrieval and a different user simulator. The simulator is part of the measurement, so this is a separate column, not a second opinion.
- Comparison source
- Artificial Analysis · Intelligence Index evaluations
- Version / tasks
- τ³-Banking (Artificial Analysis) · banking-97-tasks-gpt-5-4-mini-simulator
- Metric / unit
- pass_at_1 · percent
MCP-Atlas · Scale
Real-world multi-step tool use across 1,000 tasks, 36 MCP servers and 220 tools. The comparable lane uses Scale's April 2026 rescored protocol, all public and private tasks, a 100-tool-call budget and the exact row-specific publication date.
- Comparison source
- Scale AI Labs · MCP-Atlas
- Version / tasks
- MCP-Atlas April 2026 rescored protocol · all-1000-public-500-private-500
- Metric / unit
- pass_rate · percent
SkillsBench v1.1
Audited agent benchmark across 87 professional tasks, with supplied-skills and no-skills protocols kept separate.
- Comparison source
- SkillsBench v1.1
- Version / tasks
- v1.1 · official-87-tasks-with-skills
- Metric / unit
- pass_rate · percent
APEX-Agents
Long-horizon professional work across banking, consulting and legal environments using office tools and code execution.
- Comparison source
- Epoch AI mirror · APEX-Agents
- Version / tasks
- APEX-Agents current leaderboard · 480-professional-tasks-33-worlds
- Metric / unit
- pass_at_1 · percent
Berkeley Function Calling Leaderboard V4
Function calling, multi-turn tool use and hallucination resistance in the official BFCL V4 suite.
- Comparison source
- BFCL V4
- Version / tasks
- v4-ede5081a24bc · overall
- Metric / unit
- overall_accuracy · percent
LMArena Agent Arena
Human preference evaluation of agent outcomes. Every score retains the submitted model, disclosed thinking label, confidence interval and observation count.
- Comparison source
- LMArena Agent Arena
- Version / tasks
- 2026-09-30 · overall-ips-aggregate
- Metric / unit
- ips_score · ips
AppWorld
Interactive coding-agent benchmark across controllable applications and APIs; model, method, scaffold and task/scenario metrics remain attached to each run.
- Comparison source
- AppWorld official leaderboard
- Version / tasks
- official-c1f56015cf7c · test-challenge-all
- Metric / unit
- task_goal_completion · percent
Vending-Bench 2
Long-horizon business simulation: mean final balance in USD after a simulated year.
- Comparison source
- Epoch AI mirror · Vending-Bench 2
- Version / tasks
- Vending-Bench 2 current leaderboard · one-simulated-year-mean-final-balance
- Metric / unit
- mean_final_balance_usd · USD
GDPval-AA v2.1
Successor of GDPval-AA v2: the same 220 tasks and judges, Elo refitted and pinned to DeepSeek V4.1 Flash (max) at 1600 - not comparable with v2 values.
- Comparison source
- Artificial Analysis · Intelligence Index evaluations
- Version / tasks
- GDPval-AA v2.1 (Crowd-BT Elo, DeepSeek V4.1 Flash (max) at 1600) · 220-tasks-elo-v2-1
- Metric / unit
- elo · points
What can the data really tell us?
We look beyond which model scores higher. These analyses show where evidence supports a comparison, where measurements are missing and which conclusions would go too far.
Metrics are calculated from the same snapshot as the comparison table. The interpretation is by LLM Benchmarks; measurements come from the attributed sources.
How to use these analyses
Model selection and the availability of published tests shape every analysis. More evidence means more information to inspect, not automatically a better model. Each report states its counting unit, snapshot and limitations. Reports are recalculated with the published dataset and are not independent replications.
Read concrete model comparisons and case studies →
01Small open models: How much evidence supports a comparison?
Selecting a small open model takes more than one good score. For the selected open models below 36 billion total parameters, we examine how many model/benchmark pairs have evidence and where directly comparable evidence exists.
Broad measurement coverage helps with shortlisting. It measures available evidence, not model quality or memory requirements.
How much evidence is available?
A pair is one model and one benchmark from the full comparison table. Multiple thinking levels count only once; quantization studies do not count as base-model evidence. Comparable means at least one conflict-free result has the table's comparable status. All recorded base-model configurations are considered, not just a currently filtered view.
Each bar covers all possible pairs in its domain. Counts: comparable / documented / possible.
Our count from the public data export. Snapshot: . JSON · CSV
What does this mean for model selection?
Look for evidence in your task domain first. Many knowledge measurements do not fill a gap in coding-agent evaluation. Next, compare benchmarks shared by your candidates and check their protocols. Unknown total parameter counts are not inferred from model names and do not qualify for this under-36B analysis.
The light parts of the chart are information gaps in our dataset. A model may have been evaluated elsewhere without us having a publishable record. Model selection is also curated. The bars therefore do not establish a general advantage for open or closed models.
Reproducing the analysis
In the JSON export, select open models with a documented total parameter count above zero and below 36 billion. Include only website benchmarks. Deduplicate base-model evidence by model and benchmark; keep quantized variants separate. The dataset includes older results still eligible for publication. Also check individual measurement dates before deciding what to use.
02Same benchmark, different conditions
Two published scores can both be correct without supporting a fair difference. We count how many model/benchmark pairs have comparable evidence and explain the distinction using our evaluation lanes.
Measured in common does not automatically mean directly comparable. Version, tasks and setup must match before calculating a difference.
Documented does not always mean comparable
This analysis counts pairs across all selected models and website benchmarks. A pair is documented when at least one base-model result exists. Comparable pairs additionally require at least one conflict-free result with the corresponding table status. The remainder has supplementary or otherwise non-comparable evidence only. Thinking variants are not counted repeatedly.
Each bar covers all possible pairs in its domain. Counts: comparable / documented / possible.
Our count from the public data export. Snapshot: . JSON · CSV
Three situations where a difference misleads
SWE-bench: a model in a provider's agent and another in a uniform evaluation harness initially represent two systems. Our strict Verified comparison uses the defined Vals lane. For Pro, reported system results remain separate from the directly comparable lane.
HLE: a full task set and a text-only subset have different denominators. Tool use also changes the allowed solution methods. A percentage-point difference between these setups would not isolate model improvement.
SWE-rebench: changing task windows change which problems must be solved. An older documented window can form its own comparison cohort. Its rank belongs to that window and should not be equated with the current window.
A defensible comparison in three steps
Select two to four models and a relevant domain. Limit the view to benchmarks measured for all selected models. Then inspect comparability labels and measurement details before interpreting reference-model differences. Even a technically valid difference is not a statistical significance test.
Our conclusion: an additional provider result adds information without automatically creating another fair rank. Making this distinction visible is more useful than forcing every available score into one ranking.
03More thinking does not guarantee a better score
More computation is often equated with a better outcome. We search the published dataset for documented examples where a higher fixed thinking level has a worse score under the same disclosed evaluation setup.
A counterexample challenges the assumption of a guaranteed improvement. It does not mean less thinking is generally better.
Examples from this snapshot
We compare fixed documented levels within the same model, source, test version, task set, metric, tools, agent orchestration and evaluation. Reported sample counts must also match. We use current, fresh canonical-lane evidence without a stated conflict or contamination warning. Ambiguous multiple scores at one effort level are omitted.
Claude Fable 5.1 · CritPt
CritPt current leaderboard · Epoch AI mirror · CritPt
Evidence date basis: 2 September 2026 / 2 September 2026. Not proof of a causal thinking effect.
Claude Fable 5 · Vending 2
Vending-Bench 2 current leaderboard · Epoch AI mirror · Vending-Bench 2
Evidence date basis: 2 September 2026 / 2 September 2026. Not proof of a causal thinking effect.
Claude Opus 4.7 · ARC-AGI-1
ARC-AGI-1 current leaderboard · Epoch AI mirror · ARC-AGI-1
Evidence date basis: 2 September 2026 / 2 September 2026. Not proof of a causal thinking effect.
What these examples establish
These are deliberately selected counterexamples, not a representative sample or an estimate of frequency. Different evaluation times and undisclosed settings can still matter. Without repeated trials and an appropriate uncertainty analysis, the cause of a score difference cannot be established.
This is why our table defaults to the highest documented effort in the selected test series rather than retrospectively choosing the best score across levels. It makes the selection rule transparent. Test quality, latency and cost together for your application; the quality scores shown here alone do not form an efficiency ranking.
Who runs this site?
LLM Benchmarks is a privately operated project by Christopher Böhm. It brings published model evaluations together with their sources and limitations in one place.
A traceable selection matters more to us than a seemingly definitive leaderboard. Every comparison should lead back to its evidence.
Our contribution
We reconcile model identities and benchmark versions, separate differing test conditions, label uncertain or conflicting information and provide comparison and export tools. These guides and analyses explain those decisions. We do not claim third-party measurements as our own tests.
Responsibility and maintenance
Christopher Böhm is the operator and contact person. Data collection and technical checks are partly automated. Publication rules, source mappings and curated supplementary evidence are maintained in the project. Automation can propagate errors, which is why source links, measurement details and the correction contact are part of the service.
Funding and selection rules
The operator accepts voluntary support through PayPal. The website is also prepared for a manual AdSense advertising slot. Model inclusion, measurement selection and ranks follow the data rules described in the methodology. Source attribution documents where a result came from; it is not a product endorsement.
Contact and corrections
Found an incorrect score, a model mismatch or a missing limitation? Send the model name, benchmark, affected information and, if possible, original evidence. Suggestions for missing published evaluations are welcome too. Please do not send confidential data.