Running models locally used to mean compromising on quality for the sake of privacy. In 2026, the quality gap has narrowed enough that local-first isn't a sacrifice — it's a strategic choice. Here's what the landscape actually looks like when you run the numbers yourself.
The open-weights ecosystem has matured fast. Here's what's worth running locally right now:
| Model | Params | Min VRAM | Best For | Notes |
|---|---|---|---|---|
| Llama 3.1 8B | 8B | 8 GB | General chat, coding | Best price/performance in class |
| Mistral 7B v0.3 | 7B | 6 GB | Fast inference, instruction | Strong at Q5, great latency |
| Qwen2.5-Coder 7B | 7B | 8 GB | Code generation | Competitive with GPT-3.5 on code |
| Gemma 2 9B | 9B | 10 GB | Reasoning, summarization | Google's best small model |
| Phi-3 Medium | 14B | 12 GB | Dense reasoning | Punches above weight on benchmarks |
| Llama 3.1 70B | 70B | 40 GB* | Complex tasks, analysis | *Requires 2×24GB GPUs or heavy quant |
You don't run models at FP16 locally. You quantize. The question is how much quality you're willing to trade for speed and memory. Here's what the data says:
| Format | Size Reduction | Quality Retention | When to Use |
|---|---|---|---|
| FP16 (baseline) | 1× | 100% | Research, calibration |
| Q8_0 | ~0.55× | ~99% | When you have VRAM to spare |
| Q5_K_M | ~0.38× | ~96% | Sweet spot for most uses |
| Q4_K_M | ~0.30× | ~92% | Constrained VRAM, still solid |
| Q3_K_M | ~0.25× | ~85% | Desperation mode — noticeable degradation |
The reference implementation. Written in C/C++, supports every quantization format, runs on CPU, CUDA, Metal, Vulkan. If you're running local models, you're using llama.cpp (even if indirectly). It powers Ollama, LM Studio, and most other wrappers.
The easiest on-ramp. Wraps llama.cpp with a Docker-like model management experience. ollama run llama3.1 and you're done. Trade-off: less control over quantization and inference parameters. Best for people who want local models without thinking about GGUF formats.
For serving models at scale. PagedAttention, continuous batching, OpenAI-compatible API. If you need to serve a local model to multiple users or services, vLLM is the answer. Overkill for personal use.
Apple's framework for Metal-accelerated inference. On M-series chips, MLX often outperforms llama.cpp for Apple-specific optimizations. If you're on a Mac, try both — the gap varies by model.
# Quick benchmark: M3 Max, Llama 3.1 8B Q5_K_M
# llama.cpp (Metal): ~32 tok/s
# MLX: ~38 tok/s
# Ollama (Metal): ~30 tok/s (overhead from API layer)
Local models win on three axes: privacy, latency, and cost at scale. They lose on two: quality ceiling and convenience.
The quality gap between a local 7B model and GPT-4o or Claude 3.5 is real. On structured tasks — summarization, classification, extraction, code completion — a well-prompted 7B model at Q5 is close enough for production. On complex reasoning, creative writing, or tasks requiring broad world knowledge, the gap is significant.
# Example: Ollama setup with model pulling
ollama pull llama3.1:8b-q5_K_M
ollama pull qwen2.5-coder:7b-q5_K_M
# Serve with API compatibility
OLLAMA_HOST=0.0.0.0:11434 ollama serve
# Now hit it like OpenAI
curl http://localhost:11434/v1/chat/completions \
-d '{"model":"llama3.1:8b-q5_K_M","messages":[{"role":"user","content":"Explain quantization"}]}'
"The question isn't whether local models are good enough. It's whether they're good enough for this specific task. Answer that per-task and you'll save a fortune."
Local models in 2026 are viable for the long tail of AI tasks — the high-volume, low-stakes work that would cost a fortune at cloud API rates. Set them up right (Q5_K_M, adequate VRAM, smart routing) and they handle 80% of the workload at 0% of the per-token cost. Route the remaining 20% to the best cloud model you can afford. That's the architecture that wins.