Skip to content

What AI models can the Intel Arc B570 run?

The Intel Arc B570 pairs 10GB of VRAM with 380 GB/s of bandwidth — and bandwidth is what sets local LLM speed. Here's exactly what it can run, how fast, and at which quantization. Every number below is computed, not guessed.

At a glance

VRAM
10GB
Bandwidth
380GB/s
Runnable models
15/ 15
Largest model
70BQ2_K

Top picks for this card

Best for coding
Qwen2.5-Coder 7B
Q6_K · 50.3 tok/s · 7.2 GB
Best for reasoning
DeepSeek-R1-Distill 14B
Q3_K_M · 40.2 tok/s · 8.8 GB
Best for chat
Qwen3 14B
Q3_K_M · 40.7 tok/s · 8.7 GB

Every model on the Intel Arc B570

ModelFitBest quantMemorySpeed
Llama 3.2 3B
3.2B · Llama
PerfectFP16 / BF167.8 GB46.1 tok/s
Qwen3 4B
4B · Qwen
PerfectQ8_05.7 GB66.3 tok/s
Mistral 7B v0.3
7.2B · Mistral
PerfectQ6_K7.4 GB48.5 tok/s
Qwen2.5-Coder 7B
7B · Qwen
PerfectQ6_K7.2 GB50.3 tok/s
Llama 3.1 8B
8B · Llama
PerfectQ6_K8.1 GB44 tok/s
Qwen3 8B
8B · Qwen
PerfectQ6_K8.2 GB43.6 tok/s
Gemma 3 12B
12B · Gemma
PerfectIQ4_XS8.3 GB42.8 tok/s
Phi-4 14B
14B · Phi
PerfectQ3_K_M8.7 GB40.7 tok/s
Qwen3 14B
14B · Qwen
PerfectQ3_K_M8.7 GB40.7 tok/s
DeepSeek-R1-Distill 14B
14B · DeepSeek
PerfectQ3_K_M8.8 GB40.2 tok/s
Mistral Small 3 24B
24B · Mistral
PerfectIQ2_XXS8.3 GB43 tok/s
Gemma 3 27B
27B · Gemma
WorksIQ2_XXS9.3 GB34.9 tok/s
Qwen3 30B-A3B (MoE)
30.5B · Qwen
WorksIQ3_M16 GB38 tok/s
Qwen3 32B
32B · Qwen
WorksIQ2_XXS11 GB13.7 tok/s
Llama 3.3 70B
70B · Llama
HeavyQ2_K27 GB1.6 tok/s

Fit and speed assume a 8K context on 10GB VRAM with 32GB system RAM for overflow. Tokens/sec is single-stream decode. Click any model for its full hardware requirements.

Frequently asked

How many local LLMs can the Intel Arc B570 run?

15 of the 15 models we test fit on the Intel Arc B570's 10GB of VRAM at a sensible quantization. The largest is Llama 3.3 70B (70B) at Q2_K, generating about 1.6 tokens/sec.

How fast is the Intel Arc B570 for local AI?

Single-stream generation is memory-bandwidth bound, and the Intel Arc B570 has 380 GB/s. On an 8B model at Q4 that translates to roughly 43.6 tokens/sec — comfortably faster than reading speed.

What should I run the Intel Arc B570 with?

LM Studio or llama.cpp with the Vulkan backend.

Can the Intel Arc B570 run 70B models?

Yes — a 70B model runs, but only by offloading part of it to system RAM, which drops speed significantly. A 24-32B model is the sweet spot for staying fully in memory.

Compare other GPUs

Popular models to run

Best models by need

Run the interactive advisor
Auto-detect your exact hardware and get personalised picks, speed & memory.