# LLM Benchmarks > A sortable, source-attributed dataset of current LLM benchmark results, including distinct reasoning profiles. Snapshot generated: 2026-10-01T12:31:08Z Models: 100 Benchmarks: 54 Results: 4712 Primary table values: 2111 Use the primary table CSV for quick comparisons. Use the measurements CSV when reasoning effort, token budget, tools, scaffold or alternate protocols matter. Scores from different benchmark releases, versions, splits, metrics or units must not be compared directly. In API v1, use the explicit `directly_comparable` field; `comparison_lane=contamination-risk` is a warning and not an automatic exclusion when that field is true. Important interpretation rules: - A missing value means no exact result is available; it is never a score of zero. - Results without a trustworthy value date, or older than twelve calendar months, are excluded. - Values between three and twelve calendar months old are marked stale in the JSON/CSV and UI. - Agent benchmarks describe a model plus scaffold, tools and budget. Inspect provenance fields. - `is_primary=true` marks the stable, protocol-aware table value; other rows are retained as separate runs. - `metric_details.comparison_warning_code` is stable for machines; `comparison_warning` is its canonical English fallback, while the website localizes known codes. - Normalized effort labels preserve disclosed provider-mode ordering but are not compute-equivalent across models or providers; only a published token budget is quantitative evidence. - Model selection is an exact top 100: up to 56 Epoch ECI slots, 20 challenge slots, 2 verified recent provider releases, and 22 qualified open-weight slots. The seven exact local-user-priority Qwen checkpoints retain priority. Documented open-weight models below 36B total parameters are prioritized in the remaining open block. Admission requires independent primary evidence or the documented revision-pinned provider-card qualification for small checkpoints; supplemental provider measurements never become primary merely through admission. Any unfilled quota uses the deterministic rank-fusion fallback. Exact checkpoint identity, source evidence and metadata remain mandatory. - CSV string cells that could be interpreted as spreadsheet formulas are prefixed with an apostrophe. Ignore that documented safety prefix when reconstructing an original string, or use the lossless JSON/API contracts instead. Typed numeric CSV cells are unchanged. - Only aggregate scores and metadata are published. Benchmark questions are not redistributed. ## Data - [https://llmbenchmarks.io/api/v1/table.json](https://llmbenchmarks.io/api/v1/table.json): Compact, versioned JSON table for Home Assistant, agents and other clients. - [https://llmbenchmarks.io/api/v1/measurements.json](https://llmbenchmarks.io/api/v1/measurements.json): Versioned JSON containing every published measurement and its evidence fingerprints. - [https://llmbenchmarks.io/api/v1/changes.json](https://llmbenchmarks.io/api/v1/changes.json): Semantic changes between validated snapshots; poll timestamps alone do not create entries. - [https://llmbenchmarks.io/index.md](https://llmbenchmarks.io/index.md): Concise Markdown guide to this snapshot and its interpretation. - [https://llmbenchmarks.io/data/table.csv](https://llmbenchmarks.io/data/table.csv): Small, denormalized CSV containing only the directly comparable primary table values. - [https://llmbenchmarks.io/data/measurements.csv](https://llmbenchmarks.io/data/measurements.csv): Denormalized CSV containing every published reasoning, tool, scaffold and protocol variant. - [https://llmbenchmarks.io/data/latest.json](https://llmbenchmarks.io/data/latest.json): Complete normalized dataset with models, reasoning profiles, releases, scores and provenance. - [https://llmbenchmarks.io/data/changes.json](https://llmbenchmarks.io/data/changes.json): JSON Feed 1.1 for change subscribers. - [https://llmbenchmarks.io/rss.xml](https://llmbenchmarks.io/rss.xml): the same change feed as RSS 2.0, for readers and aggregators that expect it. - [https://llmbenchmarks.io/huggingface/README.md](https://llmbenchmarks.io/huggingface/README.md): Hugging Face Dataset Viewer compatible mirror bundle. ## Optional - [https://llmbenchmarks.io/data/latest.csv](https://llmbenchmarks.io/data/latest.csv): Stable normalized long-form CSV contract using IDs. - [https://llmbenchmarks.io/data/schema.json](https://llmbenchmarks.io/data/schema.json): JSON Schema and missing-value semantics. - [https://llmbenchmarks.io/data/source-status.json](https://llmbenchmarks.io/data/source-status.json): Collector freshness and quarantine status by registry source group; these group IDs are not row-level result source IDs. - [https://llmbenchmarks.io/openapi.json](https://llmbenchmarks.io/openapi.json): Static data endpoint description.