Skip to content

What AI models can the GeForce RTX 3090 run?

The GeForce RTX 3090 pairs 24GB of VRAM with 936 GB/s of bandwidth — and bandwidth is what sets local LLM speed. Here's exactly what it can run, how fast, and at which quantization. Every number below is computed, not guessed.

At a glance

VRAM
24GB
Bandwidth
936GB/s
Runnable models
15/ 15
Largest model
70BIQ2_XXS

Top picks for this card

Best for coding
Qwen2.5-Coder 32B
IQ4_XS · 41.9 tok/s · 19 GB
Best for reasoning
DeepSeek-R1-Distill 32B
IQ4_XS · 41.9 tok/s · 19 GB
Best for chat
Qwen3 32B
IQ4_XS · 41.9 tok/s · 19 GB

Every model on the GeForce RTX 3090

ModelFitBest quantMemorySpeed
Llama 3.2 3B
3.2B · Llama
PerfectFP16 / BF167.8 GB113.7 tok/s
Qwen3 4B
4B · Qwen
PerfectFP16 / BF169.4 GB90.8 tok/s
Mistral 7B v0.3
7.2B · Mistral
PerfectFP16 / BF1616 GB51.4 tok/s
Qwen2.5-Coder 7B
7B · Qwen
PerfectFP16 / BF1615 GB53.1 tok/s
Llama 3.1 8B
8B · Llama
PerfectFP16 / BF1618 GB46.4 tok/s
Qwen3 8B
8B · Qwen
PerfectFP16 / BF1618 GB46.3 tok/s
Gemma 3 12B
12B · Gemma
PerfectQ8_015 GB56.2 tok/s
Phi-4 14B
14B · Phi
PerfectQ8_017 GB48.9 tok/s
Qwen3 14B
14B · Qwen
PerfectQ8_017 GB48.9 tok/s
DeepSeek-R1-Distill 14B
14B · DeepSeek
PerfectQ8_017 GB48.7 tok/s
Mistral Small 3 24B
24B · Mistral
PerfectQ5_K_M19 GB42.5 tok/s
Gemma 3 27B
27B · Gemma
PerfectQ4_K_M19 GB43.6 tok/s
Qwen3 30B-A3B (MoE)
30.5B · Qwen
PerfectQ4_K_M20 GB365.7 tok/s
Qwen3 32B
32B · Qwen
PerfectIQ4_XS19 GB41.9 tok/s
Llama 3.3 70B
70B · Llama
RecommendedIQ2_XXS22 GB37.2 tok/s

Fit and speed assume a 8K context on 24GB VRAM with 32GB system RAM for overflow. Tokens/sec is single-stream decode. Click any model for its full hardware requirements.

Community-measured speeds

ModelEstimatedMeasured (median)Reports
Llama 3.1 8B46.4 tok/s118 tok/s4 (−1)
Qwen3 30B-A3B (MoE)365.7 tok/s95 tok/s1
Qwen3 14B48.9 tok/s58 tok/s1
Qwen3 32B41.9 tok/s30 tok/s1

Real generation speeds reported by the community (median, with statistical outliers trimmed), shown against our estimate to keep the model honest. Run the GeForce RTX 3090? Add your own numbers in the advisor.

Frequently asked

How many local LLMs can the GeForce RTX 3090 run?

15 of the 15 models we test fit on the GeForce RTX 3090's 24GB of VRAM at a sensible quantization. The largest is Llama 3.3 70B (70B) at IQ2_XXS, generating about 37.2 tokens/sec.

How fast is the GeForce RTX 3090 for local AI?

Single-stream generation is memory-bandwidth bound, and the GeForce RTX 3090 has 936 GB/s. On an 8B model at Q4 that translates to roughly 46.3 tokens/sec — comfortably faster than reading speed.

What should I run the GeForce RTX 3090 with?

Ollama or LM Studio (CUDA) for one-click setup, or vLLM/TabbyAPI for serving.

Can the GeForce RTX 3090 run 70B models?

Yes — a 70B model runs, but only by offloading part of it to system RAM, which drops speed significantly. A 24-32B model is the sweet spot for staying fully in memory.

Compare other GPUs

Popular models to run

Best models by need

Run the interactive advisor
Auto-detect your exact hardware and get personalised picks, speed & memory.