Skip to content

GeForce RTX 4080 SUPER vs GeForce RTX 4090 for local AI

Which is the better card for running LLMs locally? The verdict comes down to memory bandwidth and VRAM — here's the head-to-head with real, computed tokens/sec, not marketing numbers.

Verdict

For local LLMs the GeForce RTX 4090 is the stronger pick — about 317% faster on typical models thanks to its 1008 GB/s of memory bandwidth, and its 24GB of VRAM runs a 123B model.

Head to head

GeForce RTX 4080 SUPERGeForce RTX 4090
VRAM16 GB24 GB
Memory bandwidth736 GB/s1008 GB/s
FP16 compute52 TFLOPS83 TFLOPS
Largest LLM123B123B
Launch price$999$1599
Released20242022

Real LLM speed (tokens/sec, Q4_K_M)

ModelGeForce RTX 4080 SUPERGeForce RTX 4090Diff
Qwen3 8B111.4152.5+37%
Qwen3 14B65.189.1+37%
Qwen3 32B4.140.1+878%

Single-stream decode at 8K context, Q4_K_M, computed from each card’s real memory bandwidth. A model that spills to system RAM on the smaller card shows a larger gap.

Frequently asked

Is the GeForce RTX 4090 faster than the GeForce RTX 4080 SUPER for LLMs?

On average across typical models the GeForce RTX 4090 is about 317% faster. Local generation is memory-bandwidth bound, and the GeForce RTX 4090 has the higher bandwidth (1008 GB/s), which is what usually decides it.

Is upgrading from the GeForce RTX 4080 SUPER to the GeForce RTX 4090 worth it for AI?

On a 14B model at Q4 the GeForce RTX 4090 is about 37% faster (65.1 → 89.1 tok/s). It also has 8GB more VRAM, letting you run larger models.

Which has more VRAM, the GeForce RTX 4080 SUPER or GeForce RTX 4090?

The GeForce RTX 4090 has more — 24GB vs 16GB. More VRAM means larger models fit fully in memory instead of spilling to slow system RAM.

More GPU comparisons

What can these run?

Run the interactive advisor
Auto-detect your exact hardware and get personalised picks, speed & memory.