Skip to content

What AI models can the Apple M4 run?

The Apple M4 pairs 16GB of unified memory with 120 GB/s of bandwidth — and bandwidth is what sets local LLM speed. Here's exactly what it can run, how fast, and at which quantization. Every number below is computed, not guessed.

At a glance

Unified memory
16GB
Bandwidth
120GB/s
Runnable models
14/ 15
Largest model
32BQ2_K

Top picks for this card

Best for coding
Qwen2.5-Coder 7B
Q4_K_M · 21 tok/s · 5.7 GB
Best for reasoning
Qwen3 30B-A3B (MoE)
Q2_K · 82.6 tok/s · 12 GB
Best for chat
Qwen3 30B-A3B (MoE)
Q2_K · 82.6 tok/s · 12 GB

Every model on the Apple M4

ModelFitBest quantMemorySpeed
Llama 3.2 3B
3.2B · Llama
PerfectQ4_K_M3.3 GB43.1 tok/s
Qwen3 4B
4B · Qwen
PerfectQ3_K_M3.4 GB40.9 tok/s
Mistral 7B v0.3
7.2B · Mistral
RecommendedQ4_K_M5.9 GB20.2 tok/s
Qwen2.5-Coder 7B
7B · Qwen
RecommendedQ4_K_M5.7 GB21 tok/s
Llama 3.1 8B
8B · Llama
RecommendedIQ4_XS5.8 GB20.6 tok/s
Qwen3 8B
8B · Qwen
RecommendedIQ4_XS5.8 GB20.3 tok/s
Gemma 3 12B
12B · Gemma
RecommendedQ2_K5.8 GB20.3 tok/s
Phi-4 14B
14B · Phi
RecommendedQ2_K6.4 GB18.1 tok/s
Qwen3 14B
14B · Qwen
RecommendedQ2_K6.4 GB18.1 tok/s
DeepSeek-R1-Distill 14B
14B · DeepSeek
RecommendedIQ2_XXS5.5 GB21.8 tok/s
Mistral Small 3 24B
24B · Mistral
WorksIQ3_M13 GB8.2 tok/s
Gemma 3 27B
27B · Gemma
WorksQ2_K11 GB9.7 tok/s
Qwen3 30B-A3B (MoE)
30.5B · Qwen
PerfectQ2_K12 GB82.6 tok/s
Qwen3 32B
32B · Qwen
WorksQ2_K13 GB8.3 tok/s
Llama 3.3 70B
70B · Llama
Won't run46 GB tok/s

Fit and speed assume a 8K context on 16GB unified memory with 32GB system RAM for overflow. Tokens/sec is single-stream decode. Click any model for its full hardware requirements.

Community-measured speeds

ModelEstimatedMeasured (median)Reports
Llama 3.1 8B20.6 tok/s19 tok/s1

Real generation speeds reported by the community (median, with statistical outliers trimmed), shown against our estimate to keep the model honest. Run the Apple M4? Add your own numbers in the advisor.

Frequently asked

How many local LLMs can the Apple M4 run?

14 of the 15 models we test fit on the Apple M4's 16GB of unified memory at a sensible quantization. The largest is Qwen3 32B (32B) at Q2_K, generating about 8.3 tokens/sec.

How fast is the Apple M4 for local AI?

Single-stream generation is memory-bandwidth bound, and the Apple M4 has 120 GB/s. On an 8B model at Q4 that translates to roughly 20.3 tokens/sec — comfortably faster than reading speed.

What should I run the Apple M4 with?

LM Studio with the MLX backend is fastest on Apple Silicon; Ollama and llama.cpp (Metal) also work great.

Can the Apple M4 run 70B models?

Not fully in unified memory — a 70B model needs ~40GB+ at Q4. You can still run it by spilling into system RAM, but it will be slow. Stick to 14B-class models for a snappy experience.

Compare other GPUs

Popular models to run

Best models by need

Run the interactive advisor
Auto-detect your exact hardware and get personalised picks, speed & memory.