LLM Benchmarks

Source-backed measurements

Analyses & model case studies

Compare small open LLMs using concrete measurements and interpret different SWE-bench results. Original analyses with sources and explicit limitations.

Three open models below 36 billion parameters

Qwen 3.8 27B, Qwen 3.6 35B-A3B and Gemma 4 31B IT are our editorial sample: two providers, including dense and MoE architectures. This selection is not a leaderboard. The first question is whether comparable measurements exist for the same task.

Qwen 3.8 27B

Total parameters: 27 billion

Dense model

Qwen 3.6 35B-A3B

Total parameters: 35 billion

Active parameters: 3 billion

Mixture of experts (MoE)

Gemma 4 31B IT

Total parameters: 30.7 billion

Dense model

For each model and benchmark, we use the base result selected by the table rules. Quantized variants, stale or no longer confirmed evidence, conflicts and contamination warnings are excluded. Within each chart, source, test version, tasks, metric, agent, tools and judge must match. Different thinking settings remain visible; this does not establish equal compute budgets.

Scale 0–100%. Only values within the same group share the disclosed setup. We calculate no advantage between groups.

GPQA

Individual evidence – no matching comparison partner in this sample
Vals AI GPQA Diamond · GPQA Diamond v1 · Vals dual-prompt · diamond-198-zero-and-five-shot-cot
Qwen 3.8 27B88.89 %
Reasoning effort: xhigh

Original evidence: Vals AI GPQA Diamond ↗

Evidence date (source or observation):

No current base result meets these selection rules: Qwen 3.6 35B-A3B · Gemma 4 31B IT.

Benchmark guide →

SciCode

Shared disclosed evaluation setup
Epoch AI mirror · SciCode · SciCode current leaderboard · scientific-research-coding-subproblems
Qwen 3.8 27B46.6 %
Reasoning effort: xhigh

Original evidence: Epoch AI mirror · SciCode ↗

Evidence date (source or observation):

Qwen 3.6 35B-A3B1.3 %
Reasoning effort: off

Original evidence: Epoch AI mirror · SciCode ↗

Evidence date (source or observation):

Gemma 4 31B IT43.4 %
Reasoning effort: Not reported

Original evidence: Epoch AI mirror · SciCode ↗

Evidence date (source or observation):

Highest displayed point estimate within this group: Qwen 3.8 27B.

Benchmark guide →

LiveCodeBench

Individual evidence – no matching comparison partner in this sample
Vals AI · LiveCodeBench · Vals LiveCodeBench v1 · vals-livecodebench-overall
Qwen 3.8 27B84 %
Reasoning effort: xhigh

Original evidence: Vals AI · LiveCodeBench ↗

Evidence date (source or observation):

No current base result meets these selection rules: Qwen 3.6 35B-A3B · Gemma 4 31B IT.

Benchmark guide →

A higher point estimate is not proof of statistical superiority. Disclosed setups may still contain unreported differences.

GPQA is evidence about scientific problem solving. SciCode covers scientific programming; LiveCodeBench covers competitive programming. A lead on GPQA therefore does not establish which model will handle your code better. Start with your task domain and compare values only within the same group.

Our conclusion: relevant shared evidence is more useful for shortlisting than an average across these three tests. Without a shared measurement, a knowledge gap remains. Total parameter count alone establishes neither memory requirements nor the speed of a quantized local installation.

Open these three models in the comparison table →

↑ Back to topics

The same model, two SWE-bench results

Claude Opus 5 illustrates why model and benchmark names alone are insufficient. We place the independent Vals run alongside the supplementary provider report. Each result remains tied to its own source and evaluation setup.

Uniform comparison run

Claude Opus 5 · SWE-bench

97 %

Disclosed evaluation setup

  • SWE-bench Verified v1
  • verified-500-vals-mini-swe-agent
  • Vals AI mini-swe-agent
  • Bash
  • Vals AI independent run over all 500 verified tasks; uniform minimal bash-only mini-swe-agent harness

Original evidence: Vals AI SWE-bench Verified ↗

Evidence date (source or observation):

Supplementary provider report

Claude Opus 5 · SWE-bench

96 %

Disclosed evaluation setup

  • Anthropic 500-task evaluation
  • verified-500
  • unspecified agentic harness
  • not reported
  • adaptive thinking; max effort; five-trial average

Original evidence: Anthropic · Claude Opus 5 System Card ↗

Evidence date (source or observation):

The Vals record describes mini-swe-agent and Bash. The provider record describes a different or incompletely disclosed setup. Thinking and repetition details differ too. The score gap therefore cannot be attributed to a single cause; it proves neither a model improvement nor an error by either source.

To compare different models, we would start with the uniform Vals group. We use the provider report as additional evidence for the system tested there. Your own deployment still needs evaluation on your repository, with your agent and a fixed time and token budget.

↑ Back to topics

Use these examples to build your shortlist

  1. Define your task and select two to four candidates.
  2. Check shared measurements and their setups; leave gaps explicit.
  3. Test the remaining models on your own tasks for quality, latency and cost.
Methodology and further analyses →