Three open models below 36 billion parameters
Qwen 3.8 27B, Qwen 3.6 35B-A3B and Gemma 4 31B IT are our editorial sample: two providers, including dense and MoE architectures. This selection is not a leaderboard. The first question is whether comparable measurements exist for the same task.
Qwen 3.8 27B
Total parameters: 27 billion
Dense model
Qwen 3.6 35B-A3B
Total parameters: 35 billion
Active parameters: 3 billion
Mixture of experts (MoE)
Gemma 4 31B IT
Total parameters: 30.7 billion
Dense model
For each model and benchmark, we use the base result selected by the table rules. Quantized variants, stale or no longer confirmed evidence, conflicts and contamination warnings are excluded. Within each chart, source, test version, tasks, metric, agent, tools and judge must match. Different thinking settings remain visible; this does not establish equal compute budgets.
Scale 0–100%. Only values within the same group share the disclosed setup. We calculate no advantage between groups.
GPQA
Vals AI GPQA Diamond · GPQA Diamond v1 · Vals dual-prompt · diamond-198-zero-and-five-shot-cot
No current base result meets these selection rules: Qwen 3.6 35B-A3B · Gemma 4 31B IT.
Benchmark guide →SciCode
Epoch AI mirror · SciCode · SciCode current leaderboard · scientific-research-coding-subproblems
Highest displayed point estimate within this group: Qwen 3.8 27B.
LiveCodeBench
Vals AI · LiveCodeBench · Vals LiveCodeBench v1 · vals-livecodebench-overall
No current base result meets these selection rules: Qwen 3.6 35B-A3B · Gemma 4 31B IT.
Benchmark guide →A higher point estimate is not proof of statistical superiority. Disclosed setups may still contain unreported differences.
GPQA is evidence about scientific problem solving. SciCode covers scientific programming; LiveCodeBench covers competitive programming. A lead on GPQA therefore does not establish which model will handle your code better. Start with your task domain and compare values only within the same group.
Our conclusion: relevant shared evidence is more useful for shortlisting than an average across these three tests. Without a shared measurement, a knowledge gap remains. Total parameter count alone establishes neither memory requirements nor the speed of a quantized local installation.
Open these three models in the comparison table →