What AI models can the Radeon RX 7900 XTX run?
The Radeon RX 7900 XTX pairs 24GB of VRAM with 960 GB/s of bandwidth — and bandwidth is what sets local LLM speed. Here's exactly what it can run, how fast, and at which quantization. Every number below is computed, not guessed.
At a glance
Top picks for this card
Every model on the Radeon RX 7900 XTX
| Model | Fit | Best quant | Memory | Speed |
|---|---|---|---|---|
| Llama 3.2 3B 3.2B · Llama | Perfect | FP16 / BF16 | 7.8 GB | 116.6 tok/s |
| Qwen3 4B 4B · Qwen | Perfect | FP16 / BF16 | 9.4 GB | 93.2 tok/s |
| Mistral 7B v0.3 7.2B · Mistral | Perfect | FP16 / BF16 | 16 GB | 52.7 tok/s |
| Qwen2.5-Coder 7B 7B · Qwen | Perfect | FP16 / BF16 | 15 GB | 54.4 tok/s |
| Llama 3.1 8B 8B · Llama | Perfect | FP16 / BF16 | 18 GB | 47.6 tok/s |
| Qwen3 8B 8B · Qwen | Perfect | FP16 / BF16 | 18 GB | 47.5 tok/s |
| Gemma 3 12B 12B · Gemma | Perfect | Q8_0 | 15 GB | 57.7 tok/s |
| Phi-4 14B 14B · Phi | Perfect | Q8_0 | 17 GB | 50.2 tok/s |
| Qwen3 14B 14B · Qwen | Perfect | Q8_0 | 17 GB | 50.2 tok/s |
| DeepSeek-R1-Distill 14B 14B · DeepSeek | Perfect | Q8_0 | 17 GB | 49.9 tok/s |
| Mistral Small 3 24B 24B · Mistral | Perfect | Q5_K_M | 19 GB | 43.6 tok/s |
| Gemma 3 27B 27B · Gemma | Perfect | Q4_K_M | 19 GB | 44.7 tok/s |
| Qwen3 30B-A3B (MoE) 30.5B · Qwen | Perfect | Q4_K_M | 20 GB | 375.1 tok/s |
| Qwen3 32B 32B · Qwen | Perfect | IQ4_XS | 19 GB | 43 tok/s |
| Llama 3.3 70B 70B · Llama | Recommended | IQ2_XXS | 22 GB | 38.1 tok/s |
Fit and speed assume a 8K context on 24GB VRAM with 32GB system RAM for overflow. Tokens/sec is single-stream decode. Click any model for its full hardware requirements.
Community-measured speeds
| Model | Estimated | Measured (median) | Reports |
|---|---|---|---|
| Llama 3.1 8B | 47.6 tok/s | 84 tok/s | 2 |
| Qwen3 14B | 50.2 tok/s | 45 tok/s | 1 |
Real generation speeds reported by the community (median, with statistical outliers trimmed), shown against our estimate to keep the model honest. Run the Radeon RX 7900 XTX? Add your own numbers in the advisor.
Frequently asked
How many local LLMs can the Radeon RX 7900 XTX run?
15 of the 15 models we test fit on the Radeon RX 7900 XTX's 24GB of VRAM at a sensible quantization. The largest is Llama 3.3 70B (70B) at IQ2_XXS, generating about 38.1 tokens/sec.
How fast is the Radeon RX 7900 XTX for local AI?
Single-stream generation is memory-bandwidth bound, and the Radeon RX 7900 XTX has 960 GB/s. On an 8B model at Q4 that translates to roughly 47.5 tokens/sec — comfortably faster than reading speed.
What should I run the Radeon RX 7900 XTX with?
LM Studio or Ollama with the ROCm build; llama.cpp Vulkan is the reliable fallback.
Can the Radeon RX 7900 XTX run 70B models?
Yes — a 70B model runs, but only by offloading part of it to system RAM, which drops speed significantly. A 24-32B model is the sweet spot for staying fully in memory.