Skip to content

What AI models can the Apple M4 Pro run?

The Apple M4 Pro pairs 24GB of unified memory with 273 GB/s of bandwidth — and bandwidth is what sets local LLM speed. Here's exactly what it can run, how fast, and at which quantization. Every number below is computed, not guessed.

At a glance

Unified memory
24GB
Bandwidth
273GB/s
Runnable models
15/ 15
Largest model
70BIQ2_XXS

Top picks for this card

Best for coding
Qwen2.5-Coder 7B
Q5_K_M · 41.3 tok/s · 6.4 GB
Best for reasoning
Qwen3 30B-A3B (MoE)
Q4_K_M · 106.7 tok/s · 20 GB
Best for chat
Qwen3 30B-A3B (MoE)
Q4_K_M · 106.7 tok/s · 20 GB

Every model on the Apple M4 Pro

ModelFitBest quantMemorySpeed
Llama 3.2 3B
3.2B · Llama
PerfectQ8_04.8 GB59.7 tok/s
Qwen3 4B
4B · Qwen
PerfectQ8_05.7 GB47.6 tok/s
Mistral 7B v0.3
7.2B · Mistral
PerfectQ4_K_M5.9 GB46 tok/s
Qwen2.5-Coder 7B
7B · Qwen
PerfectQ5_K_M6.4 GB41.3 tok/s
Llama 3.1 8B
8B · Llama
PerfectQ4_K_M6.4 GB41.8 tok/s
Qwen3 8B
8B · Qwen
PerfectQ4_K_M6.4 GB41.3 tok/s
Gemma 3 12B
12B · Gemma
RecommendedQ6_K12 GB20.8 tok/s
Phi-4 14B
14B · Phi
RecommendedQ5_K_M12 GB20.8 tok/s
Qwen3 14B
14B · Qwen
RecommendedQ5_K_M12 GB20.8 tok/s
DeepSeek-R1-Distill 14B
14B · DeepSeek
RecommendedQ5_K_M12 GB20.7 tok/s
Mistral Small 3 24B
24B · Mistral
RecommendedIQ3_M13 GB18.6 tok/s
Gemma 3 27B
27B · Gemma
RecommendedQ2_K11 GB22 tok/s
Qwen3 30B-A3B (MoE)
30.5B · Qwen
PerfectQ4_K_M20 GB106.7 tok/s
Qwen3 32B
32B · Qwen
RecommendedQ2_K13 GB18.9 tok/s
Llama 3.3 70B
70B · Llama
WorksIQ2_XXS22 GB10.8 tok/s

Fit and speed assume a 8K context on 24GB unified memory with 32GB system RAM for overflow. Tokens/sec is single-stream decode. Click any model for its full hardware requirements.

Community-measured speeds

ModelEstimatedMeasured (median)Reports
Llama 3.1 8B41.8 tok/s30 tok/s1

Real generation speeds reported by the community (median, with statistical outliers trimmed), shown against our estimate to keep the model honest. Run the Apple M4 Pro? Add your own numbers in the advisor.

Frequently asked

How many local LLMs can the Apple M4 Pro run?

15 of the 15 models we test fit on the Apple M4 Pro's 24GB of unified memory at a sensible quantization. The largest is Llama 3.3 70B (70B) at IQ2_XXS, generating about 10.8 tokens/sec.

How fast is the Apple M4 Pro for local AI?

Single-stream generation is memory-bandwidth bound, and the Apple M4 Pro has 273 GB/s. On an 8B model at Q4 that translates to roughly 41.3 tokens/sec — comfortably faster than reading speed.

What should I run the Apple M4 Pro with?

LM Studio with the MLX backend is fastest on Apple Silicon; Ollama and llama.cpp (Metal) also work great.

Can the Apple M4 Pro run 70B models?

Yes — a 70B model runs, but only by offloading part of it to system RAM, which drops speed significantly. A 24-32B model is the sweet spot for staying fully in memory.

Compare other GPUs

Popular models to run

Best models by need

Run the interactive advisor
Auto-detect your exact hardware and get personalised picks, speed & memory.