Skip to content

Gemma 2 9B vs Llama 3.1 8B

Two open-weight models, head to head — task-by-task quality, memory needs, speed and license. Scores are 0-100 from public evaluations; speed and VRAM are computed for a common 12GB card.

Verdict

The Gemma 2 9B is the stronger all-rounder (71 vs 71 average), but pick Gemma 2 9B for multilingual and Llama 3.1 8B for agent. Licenses differ: Gemma 2 9B is Gemma, Llama 3.1 8B is Llama-3.1.

Specs

Gemma 2 9BLlama 3.1 8B
Parameters9.2B8B
Context8K128K
VRAM (Q4)7.3 GB6.4 GB
Speed on RTX 407065.7 tok/s77.2 tok/s
LicenseGemmaLlama-3.1
Released2024-062024-07

Task-by-task quality

coding
68
68
reasoning
70
70
math
66
66
chat
80
78
multilingual
82
76
agent
62
66
Average7171
Gemma 2 9BLlama 3.1 8B

Can your GPU run them?

Frequently asked

Is Gemma 2 9B or Llama 3.1 8B better?

Overall the Gemma 2 9B edges it on average quality (71 vs 71 across six task areas), but it's use-case dependent — Gemma 2 9B leads on multilingual, while Llama 3.1 8B leads on agent.

Which is better for coding, Gemma 2 9B or Llama 3.1 8B?

They're level on coding (68/100 each).

Which needs more VRAM?

Gemma 2 9B is larger (9.2B vs 8B) and needs more VRAM. At Q4 they need about 7.3 GB and 6.4 GB respectively.

Which is faster?

On a 12GB RTX 4070 at Q4, Gemma 2 9B runs at ~65.7 tok/s and Llama 3.1 8B at ~77.2 tok/s.

More model comparisons

Run the interactive advisor
Auto-detect your exact hardware and get personalised picks, speed & memory.