The best LLMs for 48GB of VRAM (2026)
48GB of VRAM (≈ a Apple M3 Max (16c)) is a sweet spot for local AI. These are the strongest models that fit in 48GB with room for a real context window — ranked by all-round quality and how fast they run.
The ranking
- 1Qwen3 30B-A3B (MoE)REASONINGQ8_0 · 34 GB · 90.8 tok/sPerfect
- 2Gemma 3 12BVISIONQ4_K_M · 9.2 GB · 40.3 tok/sPerfect
- 3Gemma 3 27BVISIONIQ4_XS · 17 GB · 21 tok/sRecommended
- 4Qwen3 14BREASONINGQ3_K_M · 8.7 GB · 42.8 tok/sPerfect
- 5Qwen3 8BREASONINGQ6_K · 8.2 GB · 45.9 tok/sPerfect
- 6Qwen3 32BREASONINGIQ3_M · 17 GB · 20.6 tok/sRecommended
- 7Gemma 2 9BQ5_K_M · 8.3 GB · 45.2 tok/sPerfect
- 8GLM-4 9BQ5_K_M · 8.4 GB · 44.3 tok/sPerfect
- 9Mistral Small 3 24BQ4_K_M · 17 GB · 21.1 tok/sRecommended
- 10Gemma 2 27BIQ4_XS · 17 GB · 20.8 tok/sRecommended
Frequently asked
What's the best LLM for 48GB of VRAM?
Qwen3 30B-A3B (MoE) at Q8_0 is our top pick — it fits 48GB comfortably while running at about 90.8 tokens/sec (34 GB in memory).
How were these ranked?
We blend each model's public evaluation quality with how well it actually runs on the reference hardware (a Apple Silicon config) — so every pick is genuinely usable, not a model you can technically load but never run at speed.
Which quantization and backend should I use?
Each pick lists its recommended quant (Q4_K_M is the usual sweet spot). Run them with Ollama or LM Studio for the easiest setup; both auto-download the right GGUF.
More best-of guides
The best local LLMs for coding
The best local LLMs for reasoning
The best local LLMs for math
The best local LLMs for general chat & writing
The best local LLMs for multilingual & translation
The best local LLMs for agents & tool use
The best local LLMs for vision & OCR
The best LLMs for 8GB of VRAM
The best LLMs for 12GB of VRAM
Popular models
Run the interactive advisor
Auto-detect your exact hardware and get personalised picks, speed & memory.