LLM Benchmarks

Source-backed measurements

Guides & methodology

Understand benchmarks, compare measurements fairly and explore the sources behind LLM Benchmarks. Methodology, a benchmark guide and analyses grounded in the data.

Where do the scores come from?

A score becomes useful when you know which model was tested and under what conditions. This guide explains how we select, check and present published results in the comparison table.

Our contribution is assembling and checking the evidence. The linked benchmark operators and model providers perform the underlying evaluations.

1. Which models and sources are included?

The selection combines models from the Epoch Capabilities Index with other demanding evaluations, checked recent provider releases and a dedicated allocation for open models. This includes smaller checkpoints that have limited representation on major leaderboards. It is a sample of the market, not a complete model directory.

We use benchmark-operator results, independent evaluations and documented provider reports. Model identity, source, test version and publication conditions must be traceable. A publicly accessible score is not automatically cleared for republication. Provider reports retain their attribution.

2. When are two scores comparable?

A benchmark name alone is not enough. We account for release, version, task window, metric, unit, tools, agent setup and evaluation procedure. Each benchmark has a defined comparison lane. Results from other setups can add useful evidence, but do not silently receive the same ranking status.

  • Rank: position within the benchmark's eligible protocol; not an overall verdict on the model.
  • ≈: supplementary evidence under different conditions. Check the measurement details first.
  • Conflict: credible sources disagree for the same documented setup. No reliable single rank.
  • Empty cell: no matching visible measurement. It means neither zero points nor a failed evaluation.

3. Which thinking level is selected?

Within the same source and test series, we prefer the highest documented fixed level: max, xhigh, high, medium, low, minimal, off. We do not retrospectively pick whichever score is highest. A valid Max run can therefore score below High. Different agents, test versions or task windows are not mixed to obtain a higher effort level.

Adaptive means thinking is enabled, but does not establish a fixed maximum budget. Undisclosed settings remain unknown. Two providers' High labels do not necessarily represent the same computation budget.

4. What do dates and source grades mean?

Evaluation date, publication, dataset version and first observation are different events. If a source does not date an individual run, we may use the first time that exact score was observed. This does not prove a recent evaluation date. Measurement details identify the date basis.

Source authority and evidence quality are evaluated separately. A+, A, B and C are our categories for sources and evidence, not statistical probabilities that a score is correct. A strong source grade does not remove differences in test conditions.

5. What the table cannot establish

Small score gaps are not established performance advantages without an appropriate uncertainty analysis. Percentage points are not relative percentage improvements. A composite index also measures something different from success on a specific task; we do not average the columns into our own overall score.

Quantization and hardware results remain supplementary studies. Different backends, GPUs or task subsets prevent isolating the effect of quantization. Weight-file size, GPU capacity and measured memory usage describe different things. The published hardware tests are not in-house laboratory measurements by LLM Benchmarks.

6. Check evidence and report an error

Open a table cell to inspect its source, setup, date and available thinking variants. Download the data as CSV or JSON to inspect them independently. To report an error, include the model, benchmark, affected score and a link to the original evidence. Changed sources are checked against the publication rules again; confirmed errors are corrected in the dataset.

What does each benchmark measure?

The useful metric depends on your task. A broad capability index, a knowledge test and a coding agent evaluation measure different things. This guide helps you interpret each kind of result.

Start with your task, check the evaluation protocol, then compare the scores.

Broad capability: ECI and composite indices

The Epoch Capabilities Index combines results from different tests into a shared capability scale. It offers an initial view across task domains. Its value is not the percentage of tasks solved. New data can change estimates even for unchanged models.

Our interpretation: use an index to shortlist models, then inspect the individual tests relevant to your work. A specialist assistant can suit your task despite a lower broad index score. Scores from different indices should not be subtracted from each other.

Knowledge and reasoning: GPQA, HLE and mathematics

GPQA contains demanding expert-written questions in biology, physics and chemistry. The table tracks its Diamond subset. Such evaluations provide evidence about problem solving under the stated setup. They do not automatically measure reliability in open-ended research.

HLE text-only and full task sets remain separate. Tool access can change the task. In mathematics, task difficulty matters: a strong score on an older or easier set does not establish performance on newer difficult tasks. For a knowledge application, we would also test your own questions against verifiable sources.

Coding: a single task or an entire agent?

A coding evaluation may test a self-contained programming problem or an agent working in an existing repository. SWE-bench evaluates real software issues. Tools, agent orchestration and the test environment also influence the result.

Our interpretation: when selecting a coding agent, compare documented system setups. A narrower coding test may be more relevant to individual functions. Verified, Pro and changing Rebench task windows are not interchangeable. Success on one set is not a success forecast for your repository.

Agents and tools: the workflow is part of the test

Terminal, tool-use and multi-turn agent evaluations require a model to act within a defined environment. Available APIs, agent setup, time limits and call limits are part of what the score means. For example, the two Banking setups remain separate in our table.

Our interpretation: these results are most useful when the test resembles your workflow. Otherwise important aspects such as permissions, recovery from errors and your own success criteria remain untested. Inspect measurement details and try a small set of representative workflows yourself.

Reading instruction and preference evaluations

IFEval tests objectively verifiable instructions, such as required output constraints. This is useful when your workflow depends on following specific requirements. It does not establish comprehensive factual accuracy or the domain quality of every answer.

Arena leaderboards reflect preferences under their comparison procedure. A preferred answer is not necessarily factually correct. Our interpretation: preference evidence can inform writing-style decisions; verifiable correctness also requires domain checks.

All benchmarks in the comparison table

The directory below links to the original pages and lists our comparison lanes. Their version and task information matters more than similar benchmark names. The grouping aids navigation; it does not weight an overall ranking.

Overview

Epoch Capabilities Index

Composite capability index fitted by Epoch AI across more than 50 benchmarks.

Comparison source
Epoch Capabilities Index
Version / tasks
Epoch ECI snapshot 2026-10-01 · official-composite-fit
Metric / unit
eci · index_points
Original benchmark page ↗
Artificial Analysis Intelligence Index

Artificial Analysis' composite Intelligence Index; one value per model and reasoning setting. The protocol identifies the methodology version.

Comparison source
Artificial Analysis · Intelligence Index evaluations
Version / tasks
Artificial Analysis Intelligence Index v4.3 · intelligence-index-v4-3
Metric / unit
index_score · points
Original benchmark page ↗
Arena Text

Human preference ranking for text models.

Comparison source
Arena Text
Version / tasks
2026-09-30 · overall-style-control
Metric / unit
arena_score · points
Original benchmark page ↗
LiveBench

Frequently refreshed objective evaluation across broad capabilities.

Comparison source
LiveBench
Version / tasks
2026-06-25 · overall-seven-category
Metric / unit
category_balanced_average · percent
Original benchmark page ↗

Knowledge and reasoning

Humanity's Last Exam · Full multimodal

Comparable values use the original 2,500-question full multimodal no-tools protocol; text-only and tool-assisted runs stay separate.

Comparison source
Scale AI Labs · Humanity's Last Exam Full
Version / tasks
HLE finalized 2,500 · Scale protocol 2025-04-03 · full-multimodal-2500-no-tools
Metric / unit
accuracy · percent
Original benchmark page ↗
Humanity's Last Exam · Text-only

Artificial Analysis' text-only subset: 2,158 of the 2,500 questions, no tools. A different question set, not a weaker run of the full protocol.

Comparison source
Artificial Analysis · Intelligence Index evaluations
Version / tasks
HLE text-only (Artificial Analysis) · text-only-2158-no-tools
Metric / unit
accuracy · percent
Original benchmark page ↗
SimpleQA Verified

Short-answer factual knowledge and abstention quality.

Comparison source
Epoch AI SimpleQA Verified
Version / tasks
Epoch independent Inspect runs · verified-1000-epoch-inspect
Metric / unit
accuracy · percent
Original benchmark page ↗
GPQA Diamond

Graduate-level science questions in physics, chemistry and biology.

Comparison source
Vals AI GPQA Diamond
Version / tasks
GPQA Diamond v1 · Vals dual-prompt · diamond-198-zero-and-five-shot-cot
Metric / unit
accuracy · percent
Original benchmark page ↗
MMLU-Pro

Broad academic multiple-choice knowledge and reasoning across fourteen subjects; Vals AI's independent run over the MMLU-Pro set.

Comparison source
Vals AI · MMLU-Pro
Version / tasks
Vals MMLU-Pro v1 · vals-mmlu-pro-overall
Metric / unit
accuracy · percent
Original benchmark page ↗
FrontierMath T1–3

Independent Epoch AI runs on 295 expert-written advanced mathematics problems; Tier 4 stays separate.

Comparison source
Epoch AI · FrontierMath Tiers 1–3 v2
Version / tasks
2.0.0 · private-285-of-295-problem-tiers-1-3-set
Metric / unit
mean_score · percent
Original benchmark page ↗
ARC-AGI-2

Abstract rule induction over 360 evaluation tasks, with pass@2 reasoning variants kept separate.

Comparison source
ARC-AGI-2
Version / tasks
v2 · semi-private-120
Metric / unit
accuracy · percent
Original benchmark page ↗
FrontierMath T4

Independent Epoch AI runs on the 41 private research-level problems from the 43-problem FrontierMath Tier 4 v2 set; the two public examples and Tiers 1–3 are not mixed into this column.

Comparison source
Epoch AI · FrontierMath Tier 4 v2
Version / tasks
2.0.0 · private-41-of-43-problem-tier-4-set
Metric / unit
mean_score · percent
Original benchmark page ↗
MMMU-Pro

Multimodal expert knowledge and reasoning.

Comparison source
Vals AI · MMMU-Pro
Version / tasks
Vals MMMU-Pro v1 · vals-mmmu-pro-overall
Metric / unit
accuracy · percent
Original benchmark page ↗
CritPt

Research-level physics reasoning across 71 open-ended challenges graded automatically.

Comparison source
Artificial Analysis · Intelligence Index evaluations
Version / tasks
CritPt (Artificial Analysis) · 70-challenges-official-grading
Metric / unit
accuracy · percent
Original benchmark page ↗
SimpleBench

Trick-question common-sense reasoning where humans still beat frontier models; average of five runs.

Comparison source
Epoch AI mirror · SimpleBench
Version / tasks
SimpleBench current leaderboard · official-set-avg-at-5
Metric / unit
accuracy_avg_at_5 · percent
Original benchmark page ↗
Chess Puzzles

Epoch AI's 100 engine-generated chess puzzles with one best move each, run and dated by Epoch.

Comparison source
Epoch AI · Chess Puzzles
Version / tasks
Epoch AI Chess Puzzles v1 · 100-engine-generated-single-best-move-puzzles
Metric / unit
mean_score · percent
Original benchmark page ↗
ARC-AGI-1

Original ARC-AGI abstraction puzzles on the semi-private set as reported by ARC Prize.

Comparison source
Epoch AI mirror · ARC-AGI-1
Version / tasks
ARC-AGI-1 current leaderboard · semi-private-evaluation-pass-at-2
Metric / unit
pass_at_2_accuracy · percent
Original benchmark page ↗
IFEval

Instruction following. Provider reports and subset studies retain their scoring and sample-set labels; they do not form a uniform cross-model ranking.

Supplementary results; no shared canonical comparison lane is specified.

Original benchmark page ↗
IFBench

Precise instruction following: 300 prompts with 344 verifiable constraints, instruction-level strict accuracy as measured by the Swallow LLM Leaderboard.

Comparison source
Swallow LLM Leaderboard
Version / tasks
swallow-evaluation-instruct · ifbench-test-300-prompts-344-instructions
Metric / unit
inst_level_strict_acc · percent
Original benchmark page ↗

Coding

SWE-bench Verified

Resolution rate on 500 verified real-world software issues; the primary comparison uses one uniform minimal agent harness.

Comparison source
Vals AI SWE-bench Verified
Version / tasks
SWE-bench Verified v1 · verified-500-vals-mini-swe-agent
Metric / unit
resolved_rate · percent
Original benchmark page ↗
SWE-rebench v2

Fresh time-windowed software issue resolution.

Comparison source
SWE-rebench
Version / tasks
rolling-v2 · 2026-05-15/2026-07-01
Metric / unit
resolved_rate · percent
Original benchmark page ↗
SWE-bench-Live · Lite

Continuously refreshed software issue resolution with agent, model and submitted reasoning label kept together.

Comparison source
SWE-bench-Live · Lite
Version / tasks
SWE-bench-Live Lite · lite-300-python
Metric / unit
resolved_rate · percent
Original benchmark page ↗
DeepSWE

Original long-horizon software engineering tasks with separate harness and reasoning-effort results.

Comparison source
DeepSWE official live leaderboard
Version / tasks
DeepSWE v1.1 live leaderboard · 113-original-long-horizon-tasks
Metric / unit
pass_at_1 · percent
Original benchmark page ↗
SWE-bench Pro Public · Scale

Scale AI Labs' separate 731-task professional software-engineering benchmark. Results keep the submitted model, agent label, confidence interval and row-specific publication date; they are never mixed with SWE-bench Verified, SWE-rebench or vendor supplemental protocols. A July 2026 OpenAI audit reported substantial task-quality concerns, so SWE-Pro should be read as one imperfect signal rather than an absolute model ranking.

Comparison source
Scale AI Labs · SWE-bench Pro Public
Version / tasks
Scale public 731-task leaderboard · public-731
Metric / unit
resolved_rate · percent
Original benchmark page ↗
CursorBench

Coding-agent evaluation with explicit reasoning levels plus cost, token and step metadata.

Comparison source
Epoch AI mirror · CursorBench
Version / tasks
CursorBench current leaderboard · official-current-suite
Metric / unit
score · percent
Original benchmark page ↗
FrontierCode 1.1 · Main

Hard mergeability-focused coding tasks; harness and reasoning effort remain attached to each score.

Comparison source
Epoch AI mirror · FrontierCode 1.1
Version / tasks
FrontierCode 1.1 · main-100-hardest-tasks
Metric / unit
mergeability_rubric_mean_at_5 · percent
Original benchmark page ↗
IOI

IOI 2024 and 2025 olympiad problems evaluated by Vals AI with olympiad graders.

Comparison source
Vals AI · IOI
Version / tasks
Vals IOI v2 · vals-ioi-2024-2026-overall
Metric / unit
score · percent
Original benchmark page ↗
LiveCodeBench

Competitive programming problems released after training cutoffs; Vals AI's independent implementation, overall pass@1.

Comparison source
Vals AI · LiveCodeBench
Version / tasks
Vals LiveCodeBench v1 · vals-livecodebench-overall
Metric / unit
pass_at_1 · percent
Original benchmark page ↗
SciCode

Scientific research coding measured by subproblem accuracy with executable Python tests.

Comparison source
Epoch AI mirror · SciCode
Version / tasks
SciCode current leaderboard · scientific-research-coding-subproblems
Metric / unit
subproblem_accuracy · percent
Original benchmark page ↗
ALE-Bench

Score-based algorithmic optimisation contests (AtCoder Heuristic); performance rating.

Comparison source
Epoch AI mirror · ALE-Bench
Version / tasks
ALE-Bench current leaderboard · atcoder-heuristic-contest-set
Metric / unit
performance_rating · points
Original benchmark page ↗
Vibe Code Bench

Building complete web applications from scratch in an agent harness, evaluated by Vals AI.

Comparison source
Vals AI · Vibe Code Bench
Version / tasks
Vals Vibe Code Bench v1.1 · vals-vibe-code-bench-overall
Metric / unit
score · percent
Original benchmark page ↗
Text Arena · Coding

Human preference rating for generated React and TypeScript web applications.

Comparison source
Epoch AI mirror · Text Arena (Coding)
Version / tasks
Text Arena Coding current pool · text-arena-coding-bradley-terry
Metric / unit
bradley_terry_rating · points
Original benchmark page ↗
WeirdML

Unusual machine-learning tasks solved end-to-end by writing and running code.

Comparison source
Epoch AI mirror · WeirdML
Version / tasks
WeirdML v2 current leaderboard · weirdml-v2-task-set
Metric / unit
accuracy · percent
Original benchmark page ↗

Agents and tools

Terminal-Bench 2.0

Terminal-Bench 2.0 leaderboard over 89 hard terminal tasks; agent harness and run date stay attached to every score.

Comparison source
Epoch AI mirror · Terminal-Bench 2.0
Version / tasks
Terminal-Bench 2.0 · official-89-task-set
Metric / unit
accuracy_mean · percent
Original benchmark page ↗
Terminal-Bench 2.1

Vals AI's independent Terminal-Bench 2.1 run with one uniform harness; vendor reports remain supplemental.

Comparison source
Vals AI · Terminal-Bench 2.1
Version / tasks
Vals Terminal-Bench 2.1 v2.1 · vals-terminal-bench-2-1-overall
Metric / unit
accuracy · percent
Original benchmark page ↗
Terminal-Bench 3.0

Official agent-plus-model leaderboard over 74 hard terminal tasks; agent, effort and trial count stay attached to every score.

Comparison source
Terminal-Bench 3.0
Version / tasks
3.0 · b936386b-4afd-4ad7-aa7f-6b52bcab9853 · official-74-task-leaderboard
Metric / unit
accuracy · percent
Original benchmark page ↗
τ³-Banking · Official harness

Knowledge-intensive multi-turn banking support tasks with retrieval, tools and a simulated user, as published on the official leaderboard.

Comparison source
τ³-Banking
Version / tasks
1.0.1 · banking_knowledge
Metric / unit
pass_1 · percent
Original benchmark page ↗
τ³-Banking · Artificial Analysis harness

Artificial Analysis' own run of the same domain: 97 tasks, BM25/grep retrieval and a different user simulator. The simulator is part of the measurement, so this is a separate column, not a second opinion.

Comparison source
Artificial Analysis · Intelligence Index evaluations
Version / tasks
τ³-Banking (Artificial Analysis) · banking-97-tasks-gpt-5-4-mini-simulator
Metric / unit
pass_at_1 · percent
Original benchmark page ↗
MCP-Atlas · Scale

Real-world multi-step tool use across 1,000 tasks, 36 MCP servers and 220 tools. The comparable lane uses Scale's April 2026 rescored protocol, all public and private tasks, a 100-tool-call budget and the exact row-specific publication date.

Comparison source
Scale AI Labs · MCP-Atlas
Version / tasks
MCP-Atlas April 2026 rescored protocol · all-1000-public-500-private-500
Metric / unit
pass_rate · percent
Original benchmark page ↗
SkillsBench v1.1

Audited agent benchmark across 87 professional tasks, with supplied-skills and no-skills protocols kept separate.

Comparison source
SkillsBench v1.1
Version / tasks
v1.1 · official-87-tasks-with-skills
Metric / unit
pass_rate · percent
Original benchmark page ↗
APEX-Agents

Long-horizon professional work across banking, consulting and legal environments using office tools and code execution.

Comparison source
Epoch AI mirror · APEX-Agents
Version / tasks
APEX-Agents current leaderboard · 480-professional-tasks-33-worlds
Metric / unit
pass_at_1 · percent
Original benchmark page ↗
Berkeley Function Calling Leaderboard V4

Function calling, multi-turn tool use and hallucination resistance in the official BFCL V4 suite.

Comparison source
BFCL V4
Version / tasks
v4-ede5081a24bc · overall
Metric / unit
overall_accuracy · percent
Original benchmark page ↗
LMArena Agent Arena

Human preference evaluation of agent outcomes. Every score retains the submitted model, disclosed thinking label, confidence interval and observation count.

Comparison source
LMArena Agent Arena
Version / tasks
2026-09-30 · overall-ips-aggregate
Metric / unit
ips_score · ips
Original benchmark page ↗
AppWorld

Interactive coding-agent benchmark across controllable applications and APIs; model, method, scaffold and task/scenario metrics remain attached to each run.

Comparison source
AppWorld official leaderboard
Version / tasks
official-c1f56015cf7c · test-challenge-all
Metric / unit
task_goal_completion · percent
Original benchmark page ↗
Vending-Bench 2

Long-horizon business simulation: mean final balance in USD after a simulated year.

Comparison source
Epoch AI mirror · Vending-Bench 2
Version / tasks
Vending-Bench 2 current leaderboard · one-simulated-year-mean-final-balance
Metric / unit
mean_final_balance_usd · USD
Original benchmark page ↗
GDPval-AA v2.1

Successor of GDPval-AA v2: the same 220 tasks and judges, Elo refitted and pinned to DeepSeek V4.1 Flash (max) at 1600 - not comparable with v2 values.

Comparison source
Artificial Analysis · Intelligence Index evaluations
Version / tasks
GDPval-AA v2.1 (Crowd-BT Elo, DeepSeek V4.1 Flash (max) at 1600) · 220-tasks-elo-v2-1
Metric / unit
elo · points
Original benchmark page ↗

What can the data really tell us?

We look beyond which model scores higher. These analyses show where evidence supports a comparison, where measurements are missing and which conclusions would go too far.

Metrics are calculated from the same snapshot as the comparison table. The interpretation is by LLM Benchmarks; measurements come from the attributed sources.

How to use these analyses

Model selection and the availability of published tests shape every analysis. More evidence means more information to inspect, not automatically a better model. Each report states its counting unit, snapshot and limitations. Reports are recalculated with the published dataset and are not independent replications.

Read concrete model comparisons and case studies →

01

Small open models: How much evidence supports a comparison?

Selecting a small open model takes more than one good score. For the selected open models below 36 billion total parameters, we examine how many model/benchmark pairs have evidence and where directly comparable evidence exists.

Broad measurement coverage helps with shortlisting. It measures available evidence, not model quality or memory requirements.

How much evidence is available?

A pair is one model and one benchmark from the full comparison table. Multiple thinking levels count only once; quantization studies do not count as base-model evidence. Comparable means at least one conflict-free result has the table's comparable status. All recorded base-model configurations are considered, not just a currently filtered view.

11selected models
171 / 506documented model/benchmark pairs
100with comparable evidence
Evidence coverage by task domain

Each bar covers all possible pairs in its domain. Counts: comparable / documented / possible.

Overview19 / 19 / 44
Knowledge and reasoning34 / 71 / 165
Coding27 / 46 / 154
Agents and tools20 / 35 / 143
Comparable evidenceOther evidence onlyNo recorded evidence

Our count from the public data export. Snapshot: . JSON · CSV

What does this mean for model selection?

Look for evidence in your task domain first. Many knowledge measurements do not fill a gap in coding-agent evaluation. Next, compare benchmarks shared by your candidates and check their protocols. Unknown total parameter counts are not inferred from model names and do not qualify for this under-36B analysis.

The light parts of the chart are information gaps in our dataset. A model may have been evaluated elsewhere without us having a publishable record. Model selection is also curated. The bars therefore do not establish a general advantage for open or closed models.

Reproducing the analysis

In the JSON export, select open models with a documented total parameter count above zero and below 36 billion. Include only website benchmarks. Deduplicate base-model evidence by model and benchmark; keep quantized variants separate. The dataset includes older results still eligible for publication. Also check individual measurement dates before deciding what to use.

02

Same benchmark, different conditions

Two published scores can both be correct without supporting a fair difference. We count how many model/benchmark pairs have comparable evidence and explain the distinction using our evaluation lanes.

Measured in common does not automatically mean directly comparable. Version, tasks and setup must match before calculating a difference.

Documented does not always mean comparable

This analysis counts pairs across all selected models and website benchmarks. A pair is documented when at least one base-model result exists. Comparable pairs additionally require at least one conflict-free result with the corresponding table status. The remainder has supplementary or otherwise non-comparable evidence only. Thinking variants are not counted repeatedly.

100selected models
2,221 / 4,600documented model/benchmark pairs
1,991with comparable evidence
Evidence coverage by task domain

Each bar covers all possible pairs in its domain. Counts: comparable / documented / possible.

Overview275 / 275 / 400
Knowledge and reasoning705 / 816 / 1,500
Coding589 / 683 / 1,400
Agents and tools422 / 447 / 1,300
Comparable evidenceOther evidence onlyNo recorded evidence

Our count from the public data export. Snapshot: . JSON · CSV

Three situations where a difference misleads

SWE-bench: a model in a provider's agent and another in a uniform evaluation harness initially represent two systems. Our strict Verified comparison uses the defined Vals lane. For Pro, reported system results remain separate from the directly comparable lane.

HLE: a full task set and a text-only subset have different denominators. Tool use also changes the allowed solution methods. A percentage-point difference between these setups would not isolate model improvement.

SWE-rebench: changing task windows change which problems must be solved. An older documented window can form its own comparison cohort. Its rank belongs to that window and should not be equated with the current window.

A defensible comparison in three steps

Select two to four models and a relevant domain. Limit the view to benchmarks measured for all selected models. Then inspect comparability labels and measurement details before interpreting reference-model differences. Even a technically valid difference is not a statistical significance test.

Our conclusion: an additional provider result adds information without automatically creating another fair rank. Making this distinction visible is more useful than forcing every available score into one ranking.

03

More thinking does not guarantee a better score

More computation is often equated with a better outcome. We search the published dataset for documented examples where a higher fixed thinking level has a worse score under the same disclosed evaluation setup.

A counterexample challenges the assumption of a guaranteed improvement. It does not mean less thinking is generally better.

Examples from this snapshot

We compare fixed documented levels within the same model, source, test version, task set, metric, tools, agent orchestration and evaluation. Reported sample counts must also match. We use current, fresh canonical-lane evidence without a stated conflict or contamination warning. Ambiguous multiple scores at one effort level are omitted.

Claude Fable 5.1 · CritPt
xhigh31.1 %
max29.7 %

CritPt current leaderboard · Epoch AI mirror · CritPt

Evidence date basis: 2 September 2026 / 2 September 2026. Not proof of a causal thinking effect.

Claude Fable 5 · Vending 2
high5680.26 USD
max4966.64 USD

Vending-Bench 2 current leaderboard · Epoch AI mirror · Vending-Bench 2

Evidence date basis: 2 September 2026 / 2 September 2026. Not proof of a causal thinking effect.

Claude Opus 4.7 · ARC-AGI-1
high93.5 %
max92.0 %

ARC-AGI-1 current leaderboard · Epoch AI mirror · ARC-AGI-1

Evidence date basis: 2 September 2026 / 2 September 2026. Not proof of a causal thinking effect.

What these examples establish

These are deliberately selected counterexamples, not a representative sample or an estimate of frequency. Different evaluation times and undisclosed settings can still matter. Without repeated trials and an appropriate uncertainty analysis, the cause of a score difference cannot be established.

This is why our table defaults to the highest documented effort in the selected test series rather than retrospectively choosing the best score across levels. It makes the selection rule transparent. Test quality, latency and cost together for your application; the quality scores shown here alone do not form an efficiency ranking.

Who runs this site?

LLM Benchmarks is a privately operated project by Christopher Böhm. It brings published model evaluations together with their sources and limitations in one place.

A traceable selection matters more to us than a seemingly definitive leaderboard. Every comparison should lead back to its evidence.

Our contribution

We reconcile model identities and benchmark versions, separate differing test conditions, label uncertain or conflicting information and provide comparison and export tools. These guides and analyses explain those decisions. We do not claim third-party measurements as our own tests.

Responsibility and maintenance

Christopher Böhm is the operator and contact person. Data collection and technical checks are partly automated. Publication rules, source mappings and curated supplementary evidence are maintained in the project. Automation can propagate errors, which is why source links, measurement details and the correction contact are part of the service.

Funding and selection rules

The operator accepts voluntary support through PayPal. The website is also prepared for a manual AdSense advertising slot. Model inclusion, measurement selection and ranks follow the data rules described in the methodology. Source attribution documents where a result came from; it is not a product endorsement.

Contact and corrections

Found an incorrect score, a model mismatch or a missing limitation? Send the model name, benchmark, affected information and, if possible, original evidence. Suggestions for missing published evaluations are welcome too. Please do not send confidential data.