Skip to content

The best local LLMs for agents & tool use (2026)

We ranked every model in our library for agents & tool use by blending public eval quality with how comfortably each one runs on typical hardware. Every pick below is genuinely runnable — no 400B models you can't load.

The ranking

  1. 1
    Qwen2.5-Coder 14B
    agent 82/100 · Q6_K · 14 GB · 46.6 tok/s
    Perfect
  2. 2
    Codestral 22B
    agent 82/100 · IQ4_XS · 14 GB · 46.9 tok/s
    Perfect
  3. 3
    Qwen3 14BREASONING
    agent 79/100 · Q6_K · 13 GB · 49.1 tok/s
    Perfect
  4. 4
    Qwen2.5-Coder 32B
    agent 86/100 · Q2_K · 13 GB · 51 tok/s
    Perfect
  5. 5
    Qwen3 32BREASONING
    agent 85/100 · Q2_K · 13 GB · 51 tok/s
    Perfect
  6. 6
    Mistral Small 3 24B
    agent 80/100 · Q3_K_M · 14 GB · 47.2 tok/s
    Perfect
  7. 7
    Qwen2.5 14B
    agent 76/100 · Q6_K · 14 GB · 46.6 tok/s
    Perfect
  8. 8
    Gemma 3 27BVISION
    agent 79/100 · IQ3_M · 15 GB · 44.2 tok/s
    Perfect
  9. 9
    Qwen3 30B-A3B (MoE)REASONING
    agent 82/100 · Q2_K · 12 GB · 506.6 tok/s
    Perfect
  10. 10
    DeepSeek-R1-Distill 32BREASONING
    agent 82/100 · Q2_K · 13 GB · 51 tok/s
    Perfect

Frequently asked

What's the best local model for agent?

Qwen2.5-Coder 14B at Q6_K is our top pick — it leads on agent while running at about 46.6 tokens/sec (14 GB in memory).

How were these ranked?

We blend each model's public evaluation quality with how well it actually runs on the reference hardware (a GeForce RTX 4080 SUPER) — so every pick is genuinely usable, not a model you can technically load but never run at speed.

Which quantization and backend should I use?

Each pick lists its recommended quant (Q4_K_M is the usual sweet spot). Run them with Ollama or LM Studio for the easiest setup; both auto-download the right GGUF.

More best-of guides

Popular models

Run the interactive advisor
Auto-detect your exact hardware and get personalised picks, speed & memory.