# LLM Benchmarks

LLM Benchmarks compares 100 current language models across 54 source-attributed benchmarks. The snapshot was generated at 2026-10-01T12:31:08Z and contains 4712 published measurements, including separate reasoning, tool, scaffold and protocol variants where sources disclose them.

## Fastest way to answer questions

1. Read [https://llmbenchmarks.io/api/v1/table.json](https://llmbenchmarks.io/api/v1/table.json) for the compact, versioned canonical API matrix across all published benchmarks. The single-page website may intentionally omit selected columns for readability; the API does not.
2. Read [https://llmbenchmarks.io/api/v1/measurements.json](https://llmbenchmarks.io/api/v1/measurements.json) when the question asks about thinking effort, token budget, tools, agent scaffold or alternative protocols.
3. Use [https://llmbenchmarks.io/data/table.csv](https://llmbenchmarks.io/data/table.csv) and [https://llmbenchmarks.io/data/measurements.csv](https://llmbenchmarks.io/data/measurements.csv) for compact tabular processing, or [https://llmbenchmarks.io/data/latest.json](https://llmbenchmarks.io/data/latest.json) for the complete normalized snapshot.
4. Read [https://llmbenchmarks.io/api/v1/changes.json](https://llmbenchmarks.io/api/v1/changes.json) to discover material model, score, thinking and canonical-table changes since prior validated snapshots.

## Interpretation rules

- For direct score comparisons, require `directly_comparable=true` and keep the same benchmark ID, `benchmark_release_id`, `benchmark_version`, `split_or_window`, metric and unit. In API v1 measurements these benchmark fields are nested below `benchmark`.
- `comparison_lane` is a protocol or warning classification, not a universal join key. In particular, a current SWE-rebench row with `comparison_lane=contamination-risk` remains directly comparable to clean rows for the same canonical release and window when `directly_comparable=true`; retain and disclose the contamination warning.
- `is_primary=true` identifies the deterministic, protocol-aware value used in the canonical API matrix. The website may omit selected benchmark columns without changing this flag.
- `higher_is_better` defines the score direction; never infer it from the benchmark name.
- `source_stated_rank` is copied from the exact source and is not the table's dynamically recomputed rank; confidence bounds use the same `unit` as the score.
- A missing row means that no eligible exact measurement is available. It never means a score of zero.
- `freshness_at` is the evidence-backed date used for age calculations. Values over three calendar months are marked `stale`; values over twelve months or without a trustworthy date are excluded.
- `reasoning_effort=unknown` means the source did not disclose a safely normalizable thinking level. Do not infer one from a model name or score.
- Reasoning-effort labels preserve disclosed provider-mode ordering but are not equivalent compute budgets across models or providers. Only compare quantitative reasoning budgets when `reasoning_token_budget` is published.
- Coding and agent scores may depend on `tools_allowed`, `scaffold`, `judge`, sample count and run count. Inspect those fields before comparing systems.
- `source_url`, evidence status and source ranks identify the source and evidence quality. For exact provenance, use `content_hash` and `source_locator` where available; a source URL may point to a changing leaderboard.
- `measurement_id` in the agent-oriented CSV files is a deterministic, evidence-bound identifier for one published measurement. It stays stable while its protocol, score, source evidence and normalized configuration stay unchanged.
- Public CSV files are spreadsheet-safe projections: formula-like string cells receive one leading apostrophe after detection through BOM, whitespace and control prefixes. JSON/API values remain lossless and unchanged; real numeric values, including negatives, remain numeric.
- `source-status.json` reports collector health by registry source group and is intentionally not a row-level join table. For result provenance, use each measurement's `source_id`, `source_name` and `source_url`.

## Machine contracts

- [API v1 metadata and semantics](https://llmbenchmarks.io/api/v1/meta.json)
- [API v1 models](https://llmbenchmarks.io/api/v1/models.json)
- [API v1 benchmarks](https://llmbenchmarks.io/api/v1/benchmarks.json)
- [API v1 table](https://llmbenchmarks.io/api/v1/table.json)
- [API v1 measurements](https://llmbenchmarks.io/api/v1/measurements.json)
- [API v1 changes](https://llmbenchmarks.io/api/v1/changes.json)
- [JSON Feed changes](https://llmbenchmarks.io/data/changes.json)
- [RSS changes](https://llmbenchmarks.io/rss.xml)
- [Dataset JSON Schema](https://llmbenchmarks.io/data/schema.json)
- [OpenAPI description](https://llmbenchmarks.io/openapi.json)
- [Complete normalized CSV](https://llmbenchmarks.io/data/latest.csv)
- [Source health](https://llmbenchmarks.io/data/source-status.json)
- [LLM discovery file](https://llmbenchmarks.io/llms.txt)

Only aggregate scores and metadata are published; benchmark questions are not redistributed. Database compilation is CC BY 4.0, while upstream terms continue to apply per result.
