Skip to content

What AI models can the Apple M3 Ultra run?

The Apple M3 Ultra pairs 96GB of unified memory with 819 GB/s of bandwidth — and bandwidth is what sets local LLM speed. Here's exactly what it can run, how fast, and at which quantization. Every number below is computed, not guessed.

At a glance

Unified memory
96GB
Bandwidth
819GB/s
Runnable models
15/ 15
Largest model
70BIQ3_M

Top picks for this card

Best for coding
Codestral 22B
Q5_K_M · 40 tok/s · 18 GB
Best for reasoning
DeepSeek-R1-Distill 32B
IQ3_M · 42.1 tok/s · 17 GB
Best for chat
Qwen3 30B-A3B (MoE)
FP16 / BF16 · 100.1 tok/s · 63 GB

Every model on the Apple M3 Ultra

ModelFitBest quantMemorySpeed
Llama 3.2 3B
3.2B · Llama
PerfectFP16 / BF167.8 GB99.5 tok/s
Qwen3 4B
4B · Qwen
PerfectFP16 / BF169.4 GB79.5 tok/s
Mistral 7B v0.3
7.2B · Mistral
PerfectFP16 / BF1616 GB45 tok/s
Qwen2.5-Coder 7B
7B · Qwen
PerfectFP16 / BF1615 GB46.4 tok/s
Llama 3.1 8B
8B · Llama
PerfectFP16 / BF1618 GB40.6 tok/s
Qwen3 8B
8B · Qwen
PerfectFP16 / BF1618 GB40.5 tok/s
Gemma 3 12B
12B · Gemma
PerfectQ8_015 GB49.2 tok/s
Phi-4 14B
14B · Phi
PerfectQ8_017 GB42.8 tok/s
Qwen3 14B
14B · Qwen
PerfectQ8_017 GB42.8 tok/s
DeepSeek-R1-Distill 14B
14B · DeepSeek
PerfectQ8_017 GB42.6 tok/s
Mistral Small 3 24B
24B · Mistral
PerfectQ4_K_M17 GB43.2 tok/s
Gemma 3 27B
27B · Gemma
PerfectIQ4_XS17 GB42.9 tok/s
Qwen3 30B-A3B (MoE)
30.5B · Qwen
PerfectFP16 / BF1663 GB100.1 tok/s
Qwen3 32B
32B · Qwen
PerfectIQ3_M17 GB42.1 tok/s
Llama 3.3 70B
70B · Llama
RecommendedIQ3_M36 GB19.4 tok/s

Fit and speed assume a 8K context on 96GB unified memory with 32GB system RAM for overflow. Tokens/sec is single-stream decode. Click any model for its full hardware requirements.

Community-measured speeds

ModelEstimatedMeasured (median)Reports
Qwen3 32B42.1 tok/s32 tok/s1
Llama 3.3 70B19.4 tok/s14 tok/s1

Real generation speeds reported by the community (median, with statistical outliers trimmed), shown against our estimate to keep the model honest. Run the Apple M3 Ultra? Add your own numbers in the advisor.

Frequently asked

How many local LLMs can the Apple M3 Ultra run?

15 of the 15 models we test fit on the Apple M3 Ultra's 96GB of unified memory at a sensible quantization. The largest is Llama 3.3 70B (70B) at IQ3_M, generating about 19.4 tokens/sec.

How fast is the Apple M3 Ultra for local AI?

Single-stream generation is memory-bandwidth bound, and the Apple M3 Ultra has 819 GB/s. On an 8B model at Q4 that translates to roughly 40.5 tokens/sec — comfortably faster than reading speed.

What should I run the Apple M3 Ultra with?

LM Studio with the MLX backend is fastest on Apple Silicon; Ollama and llama.cpp (Metal) also work great.

Can the Apple M3 Ultra run 70B models?

Yes — a 70B model runs, but it fits comfortably in memory.

Compare other GPUs

Popular models to run

Best models by need

Run the interactive advisor
Auto-detect your exact hardware and get personalised picks, speed & memory.