Skip to content

What AI models can the Apple M2 run?

The Apple M2 pairs 8GB of unified memory with 100 GB/s of bandwidth — and bandwidth is what sets local LLM speed. Here's exactly what it can run, how fast, and at which quantization. Every number below is computed, not guessed.

At a glance

Unified memory
8GB
Bandwidth
100GB/s
Runnable models
10/ 15
Largest model
14BIQ2_XXS

Top picks for this card

Best for coding
Qwen2.5-Coder 7B
IQ4_XS · 19.6 tok/s · 5.2 GB
Best for reasoning
DeepSeek-R1-Distill 7B
IQ4_XS · 19.6 tok/s · 5.2 GB
Best for chat
Qwen3 8B
IQ3_M · 19.3 tok/s · 5.3 GB

Every model on the Apple M2

ModelFitBest quantMemorySpeed
Llama 3.2 3B
3.2B · Llama
PerfectIQ4_XS3.0 GB40 tok/s
Qwen3 4B
4B · Qwen
RecommendedQ6_K4.7 GB22 tok/s
Mistral 7B v0.3
7.2B · Mistral
RecommendedIQ4_XS5.3 GB18.9 tok/s
Qwen2.5-Coder 7B
7B · Qwen
RecommendedIQ4_XS5.2 GB19.6 tok/s
Llama 3.1 8B
8B · Llama
RecommendedIQ3_M5.2 GB19.6 tok/s
Qwen3 8B
8B · Qwen
RecommendedIQ3_M5.3 GB19.3 tok/s
Gemma 3 12B
12B · Gemma
RecommendedIQ2_XXS5.0 GB20.5 tok/s
Phi-4 14B
14B · Phi
RecommendedIQ2_XXS5.4 GB18.5 tok/s
Qwen3 14B
14B · Qwen
RecommendedIQ2_XXS5.4 GB18.5 tok/s
DeepSeek-R1-Distill 14B
14B · DeepSeek
RecommendedIQ2_XXS5.5 GB18.2 tok/s
Mistral Small 3 24B
24B · Mistral
Won't run17 GB tok/s
Gemma 3 27B
27B · Gemma
Won't run19 GB tok/s
Qwen3 30B-A3B (MoE)
30.5B · Qwen
Won't run20 GB tok/s
Qwen3 32B
32B · Qwen
Won't run22 GB tok/s
Llama 3.3 70B
70B · Llama
Won't run46 GB tok/s

Fit and speed assume a 8K context on 8GB unified memory with 32GB system RAM for overflow. Tokens/sec is single-stream decode. Click any model for its full hardware requirements.

Frequently asked

How many local LLMs can the Apple M2 run?

10 of the 15 models we test fit on the Apple M2's 8GB of unified memory at a sensible quantization. The largest is Phi-4 14B (14B) at IQ2_XXS, generating about 18.5 tokens/sec.

How fast is the Apple M2 for local AI?

Single-stream generation is memory-bandwidth bound, and the Apple M2 has 100 GB/s. On an 8B model at Q4 that translates to roughly 19.3 tokens/sec — comfortably faster than reading speed.

What should I run the Apple M2 with?

LM Studio with the MLX backend is fastest on Apple Silicon; Ollama and llama.cpp (Metal) also work great.

Can the Apple M2 run 70B models?

Not fully in unified memory — a 70B model needs ~40GB+ at Q4. You can still run it by spilling into system RAM, but it will be slow. Stick to 8B-class models for a snappy experience.

Compare other GPUs

Popular models to run

Best models by need

Run the interactive advisor
Auto-detect your exact hardware and get personalised picks, speed & memory.