The best local LLMs for general chat & writing (2026)
We ranked every model in our library for general chat & writing by blending public eval quality with how comfortably each one runs on typical hardware. Every pick below is genuinely runnable — no 400B models you can't load.
The ranking
- 1Qwen3 8BREASONINGchat 80/100 · Q4_K_M · 6.4 GB · 41.2 tok/sPerfect
- 2Llama 3.1 8Bchat 78/100 · Q4_K_M · 6.4 GB · 41.7 tok/sPerfect
- 3Gemma 2 9Bchat 80/100 · Q3_K_M · 6.2 GB · 42.6 tok/sPerfect
- 4GLM-4 9Bchat 80/100 · Q3_K_M · 6.3 GB · 41.8 tok/sPerfect
- 5Qwen3 14BREASONINGchat 84/100 · Q2_K · 6.4 GB · 41.1 tok/sPerfect
- 6Command R7Bchat 76/100 · Q5_K_M · 6.4 GB · 41.2 tok/sPerfect
- 7Qwen2.5 7Bchat 76/100 · Q4_K_M · 6.0 GB · 44.2 tok/sPerfect
- 8Aya Expanse 8Bchat 76/100 · Q4_K_M · 6.4 GB · 41.7 tok/sPerfect
- 9Yi 1.5 9Bchat 76/100 · IQ4_XS · 6.0 GB · 44.2 tok/sPerfect
- 10Gemma 3 12BVISIONchat 82/100 · Q2_K · 5.8 GB · 46 tok/sPerfect
Frequently asked
What's the best local model for chat?
Qwen3 8B at Q4_K_M is our top pick — it leads on chat while running at about 41.2 tokens/sec (6.4 GB in memory).
How were these ranked?
We blend each model's public evaluation quality with how well it actually runs on the reference hardware (a GeForce RTX 4060) — so every pick is genuinely usable, not a model you can technically load but never run at speed.
Which quantization and backend should I use?
Each pick lists its recommended quant (Q4_K_M is the usual sweet spot). Run them with Ollama or LM Studio for the easiest setup; both auto-download the right GGUF.
More best-of guides
The best local LLMs for coding
The best local LLMs for reasoning
The best local LLMs for math
The best local LLMs for multilingual & translation
The best local LLMs for agents & tool use
The best local LLMs for vision & OCR
The best LLMs for 8GB of VRAM
The best LLMs for 12GB of VRAM
The best LLMs for 16GB of VRAM
Popular models
Run the interactive advisor
Auto-detect your exact hardware and get personalised picks, speed & memory.