Qwen2.5 72B vs Llama 3.3 70B
Two open-weight models, head to head — task-by-task quality, memory needs, speed and license. Scores are 0-100 from public evaluations; speed and VRAM are computed for a common 12GB card.
Verdict
The Qwen2.5 72B is the stronger all-rounder (89 vs 86 average), but pick Qwen2.5 72B for coding and Llama 3.3 70B for chat. Licenses differ: Qwen2.5 72B is Qwen, Llama 3.3 70B is Llama-3.3.
Specs
| Qwen2.5 72B | Llama 3.3 70B | |
|---|---|---|
| Parameters | 72.7B | 70B |
| Context | 128K | 128K |
| VRAM (Q4) | 48 GB | 46 GB |
| Speed on RTX 4070 | won't fit | won't fit |
| License | Qwen | Llama-3.3 |
| Released | 2024-09 | 2024-12 |
Task-by-task quality
coding
88
84
reasoning
89
87
math
88
85
chat
90
89
multilingual
90
88
agent
86
84
Average8986
Qwen2.5 72BLlama 3.3 70B
Can your GPU run them?
Frequently asked
Is Qwen2.5 72B or Llama 3.3 70B better?
Overall the Qwen2.5 72B edges it on average quality (89 vs 86 across six task areas), but it's use-case dependent — Qwen2.5 72B leads on coding, while Llama 3.3 70B leads on chat.
Which is better for coding, Qwen2.5 72B or Llama 3.3 70B?
Qwen2.5 72B — it scores 88/100 for coding vs 84/100.
Which needs more VRAM?
Qwen2.5 72B is larger (72.7B vs 70B) and needs more VRAM. At Q4 they need about 48 GB and 46 GB respectively.
Which is faster?
On a 12GB RTX 4070 at Q4, Qwen2.5 72B runs at ~— tok/s and Llama 3.3 70B at ~— tok/s.
More model comparisons
Run the interactive advisor
Auto-detect your exact hardware and get personalised picks, speed & memory.