Skip to content

What AI models can the RTX A6000 run?

The RTX A6000 pairs 48GB of VRAM with 768 GB/s of bandwidth — and bandwidth is what sets local LLM speed. Here's exactly what it can run, how fast, and at which quantization. Every number below is computed, not guessed.

At a glance

VRAM
48GB
Bandwidth
768GB/s
Runnable models
15/ 15
Largest model
70BIQ3_M

Top picks for this card

Best for coding
Codestral 22B
Q4_K_M · 43.5 tok/s · 15 GB
Best for reasoning
DeepSeek-R1-Distill 14B
Q6_K · 50.9 tok/s · 13 GB
Best for chat
Qwen3 30B-A3B (MoE)
Q8_0 · 174.3 tok/s · 34 GB

Every model on the RTX A6000

ModelFitBest quantMemorySpeed
Llama 3.2 3B
3.2B · Llama
PerfectFP16 / BF167.8 GB93.3 tok/s
Qwen3 4B
4B · Qwen
PerfectFP16 / BF169.4 GB74.5 tok/s
Mistral 7B v0.3
7.2B · Mistral
PerfectFP16 / BF1616 GB42.2 tok/s
Qwen2.5-Coder 7B
7B · Qwen
PerfectFP16 / BF1615 GB43.6 tok/s
Llama 3.1 8B
8B · Llama
PerfectQ8_010 GB69.8 tok/s
Qwen3 8B
8B · Qwen
PerfectQ8_010 GB69.3 tok/s
Gemma 3 12B
12B · Gemma
PerfectQ8_015 GB46.1 tok/s
Phi-4 14B
14B · Phi
PerfectQ8_017 GB40.1 tok/s
Qwen3 14B
14B · Qwen
PerfectQ8_017 GB40.1 tok/s
DeepSeek-R1-Distill 14B
14B · DeepSeek
PerfectQ6_K13 GB50.9 tok/s
Mistral Small 3 24B
24B · Mistral
PerfectQ4_K_M17 GB40.5 tok/s
Gemma 3 27B
27B · Gemma
PerfectIQ4_XS17 GB40.2 tok/s
Qwen3 30B-A3B (MoE)
30.5B · Qwen
PerfectQ8_034 GB174.3 tok/s
Qwen3 32B
32B · Qwen
RecommendedQ6_K29 GB22.9 tok/s
Llama 3.3 70B
70B · Llama
RecommendedIQ3_M36 GB18.2 tok/s

Fit and speed assume a 8K context on 48GB VRAM with 32GB system RAM for overflow. Tokens/sec is single-stream decode. Click any model for its full hardware requirements.

Frequently asked

How many local LLMs can the RTX A6000 run?

15 of the 15 models we test fit on the RTX A6000's 48GB of VRAM at a sensible quantization. The largest is Llama 3.3 70B (70B) at IQ3_M, generating about 18.2 tokens/sec.

How fast is the RTX A6000 for local AI?

Single-stream generation is memory-bandwidth bound, and the RTX A6000 has 768 GB/s. On an 8B model at Q4 that translates to roughly 69.3 tokens/sec — comfortably faster than reading speed.

What should I run the RTX A6000 with?

Ollama or LM Studio (CUDA) for one-click setup, or vLLM/TabbyAPI for serving.

Can the RTX A6000 run 70B models?

Yes — a 70B model runs, but it fits comfortably in memory.

Compare other GPUs

Popular models to run

Best models by need

Run the interactive advisor
Auto-detect your exact hardware and get personalised picks, speed & memory.