Skip to content

The best local LLMs for multilingual & translation (2026)

We ranked every model in our library for multilingual & translation by blending public eval quality with how comfortably each one runs on typical hardware. Every pick below is genuinely runnable — no 400B models you can't load.

The ranking

  1. 1
    Aya Expanse 8B
    multilingual 90/100 · Q8_0 · 10 GB · 45.8 tok/s
    Perfect
  2. 2
    Qwen3 14BREASONING
    multilingual 88/100 · Q4_K_M · 10 GB · 44.6 tok/s
    Perfect
  3. 3
    Gemma 3 12BVISION
    multilingual 86/100 · Q5_K_M · 10 GB · 43.9 tok/s
    Perfect
  4. 4
    Qwen3 8BREASONING
    multilingual 84/100 · Q8_0 · 10 GB · 45.5 tok/s
    Perfect
  5. 5
    GLM-4 9B
    multilingual 84/100 · Q6_K · 9.4 GB · 48.9 tok/s
    Perfect
  6. 6
    Aya Expanse 32B
    multilingual 94/100 · IQ2_XXS · 11 GB · 43.3 tok/s
    Perfect
  7. 7
    Qwen2.5 14B
    multilingual 84/100 · Q4_K_M · 11 GB · 42.3 tok/s
    Perfect
  8. 8
    Command R7B
    multilingual 82/100 · Q8_0 · 8.9 GB · 52.3 tok/s
    Perfect
  9. 9
    Gemma 2 9B
    multilingual 82/100 · Q6_K · 9.3 GB · 49.9 tok/s
    Perfect
  10. 10
    Mistral Nemo 12B
    multilingual 82/100 · Q5_K_M · 10 GB · 44.3 tok/s
    Perfect

Frequently asked

What's the best local model for multilingual?

Aya Expanse 8B at Q8_0 is our top pick — it leads on multilingual while running at about 45.8 tokens/sec (10 GB in memory).

How were these ranked?

We blend each model's public evaluation quality with how well it actually runs on the reference hardware (a GeForce RTX 4070) — so every pick is genuinely usable, not a model you can technically load but never run at speed.

Which quantization and backend should I use?

Each pick lists its recommended quant (Q4_K_M is the usual sweet spot). Run them with Ollama or LM Studio for the easiest setup; both auto-download the right GGUF.

More best-of guides

Popular models

Run the interactive advisor
Auto-detect your exact hardware and get personalised picks, speed & memory.