Skip to content

What AI models can the Radeon RX 7900 XT run?

The Radeon RX 7900 XT pairs 20GB of VRAM with 800 GB/s of bandwidth — and bandwidth is what sets local LLM speed. Here's exactly what it can run, how fast, and at which quantization. Every number below is computed, not guessed.

At a glance

VRAM
20GB
Bandwidth
800GB/s
Runnable models
15/ 15
Largest model
70BIQ4_XS

Top picks for this card

Best for coding
Codestral 22B
Q4_K_M · 45.3 tok/s · 15 GB
Best for reasoning
DeepSeek-R1-Distill 32B
IQ3_M · 41.1 tok/s · 17 GB
Best for chat
Gemma 3 27B
IQ4_XS · 41.9 tok/s · 17 GB

Every model on the Radeon RX 7900 XT

ModelFitBest quantMemorySpeed
Llama 3.2 3B
3.2B · Llama
PerfectFP16 / BF167.8 GB97.2 tok/s
Qwen3 4B
4B · Qwen
PerfectFP16 / BF169.4 GB77.6 tok/s
Mistral 7B v0.3
7.2B · Mistral
PerfectFP16 / BF1616 GB44 tok/s
Qwen2.5-Coder 7B
7B · Qwen
PerfectFP16 / BF1615 GB45.4 tok/s
Llama 3.1 8B
8B · Llama
PerfectQ8_010 GB72.7 tok/s
Qwen3 8B
8B · Qwen
PerfectQ8_010 GB72.2 tok/s
Gemma 3 12B
12B · Gemma
PerfectQ8_015 GB48.1 tok/s
Phi-4 14B
14B · Phi
PerfectQ8_017 GB41.8 tok/s
Qwen3 14B
14B · Qwen
PerfectQ8_017 GB41.8 tok/s
DeepSeek-R1-Distill 14B
14B · DeepSeek
PerfectQ8_017 GB41.6 tok/s
Mistral Small 3 24B
24B · Mistral
PerfectQ4_K_M17 GB42.2 tok/s
Gemma 3 27B
27B · Gemma
PerfectIQ4_XS17 GB41.9 tok/s
Qwen3 30B-A3B (MoE)
30.5B · Qwen
PerfectIQ4_XS18 GB352.8 tok/s
Qwen3 32B
32B · Qwen
PerfectIQ3_M17 GB41.1 tok/s
Llama 3.3 70B
70B · Llama
HeavyIQ4_XS41 GB1.3 tok/s

Fit and speed assume a 8K context on 20GB VRAM with 32GB system RAM for overflow. Tokens/sec is single-stream decode. Click any model for its full hardware requirements.

Frequently asked

How many local LLMs can the Radeon RX 7900 XT run?

15 of the 15 models we test fit on the Radeon RX 7900 XT's 20GB of VRAM at a sensible quantization. The largest is Llama 3.3 70B (70B) at IQ4_XS, generating about 1.3 tokens/sec.

How fast is the Radeon RX 7900 XT for local AI?

Single-stream generation is memory-bandwidth bound, and the Radeon RX 7900 XT has 800 GB/s. On an 8B model at Q4 that translates to roughly 72.2 tokens/sec — comfortably faster than reading speed.

What should I run the Radeon RX 7900 XT with?

LM Studio or Ollama with the ROCm build; llama.cpp Vulkan is the reliable fallback.

Can the Radeon RX 7900 XT run 70B models?

Yes — a 70B model runs, but only by offloading part of it to system RAM, which drops speed significantly. A 24-32B model is the sweet spot for staying fully in memory.

Compare other GPUs

Popular models to run

Best models by need

Run the interactive advisor
Auto-detect your exact hardware and get personalised picks, speed & memory.