← Back to Blog

Local Models Deep Dive —

📅 🏷 Essay Technical
Essay Technical

Running models locally used to mean compromising on quality for the sake of privacy. In 2026, the quality gap has narrowed enough that local-first isn't a sacrifice — it's a strategic choice. Here's what the landscape actually looks like when you run the numbers yourself.

8GB
VRAM floor for usable 7B
~85%
of cloud quality at Q4_K_M
15-30
tok/s on consumer GPU (7B)
$0
per-token cost at any scale

The Model Landscape

The open-weights ecosystem has matured fast. Here's what's worth running locally right now:

ModelParamsMin VRAMBest ForNotes
Llama 3.1 8B8B8 GBGeneral chat, codingBest price/performance in class
Mistral 7B v0.37B6 GBFast inference, instructionStrong at Q5, great latency
Qwen2.5-Coder 7B7B8 GBCode generationCompetitive with GPT-3.5 on code
Gemma 2 9B9B10 GBReasoning, summarizationGoogle's best small model
Phi-3 Medium14B12 GBDense reasoningPunches above weight on benchmarks
Llama 3.1 70B70B40 GB*Complex tasks, analysis*Requires 2×24GB GPUs or heavy quant

Quantization: The Art of Shrinking

You don't run models at FP16 locally. You quantize. The question is how much quality you're willing to trade for speed and memory. Here's what the data says:

FormatSize ReductionQuality RetentionWhen to Use
FP16 (baseline)1×100%Research, calibration
Q8_0~0.55×~99%When you have VRAM to spare
Q5_K_M~0.38×~96%Sweet spot for most uses
Q4_K_M~0.30×~92%Constrained VRAM, still solid
Q3_K_M~0.25×~85%Desperation mode — noticeable degradation
Recommendation: Default to Q5_K_M. It's the point where quality loss becomes nearly imperceptible for most tasks while cutting memory nearly in half. Drop to Q4_K_M only if you're VRAM-constrained.

Runtime Choices

llama.cpp

The reference implementation. Written in C/C++, supports every quantization format, runs on CPU, CUDA, Metal, Vulkan. If you're running local models, you're using llama.cpp (even if indirectly). It powers Ollama, LM Studio, and most other wrappers.

Ollama

The easiest on-ramp. Wraps llama.cpp with a Docker-like model management experience. ollama run llama3.1 and you're done. Trade-off: less control over quantization and inference parameters. Best for people who want local models without thinking about GGUF formats.

vLLM

For serving models at scale. PagedAttention, continuous batching, OpenAI-compatible API. If you need to serve a local model to multiple users or services, vLLM is the answer. Overkill for personal use.

MLX (Apple Silicon)

Apple's framework for Metal-accelerated inference. On M-series chips, MLX often outperforms llama.cpp for Apple-specific optimizations. If you're on a Mac, try both — the gap varies by model.

# Quick benchmark: M3 Max, Llama 3.1 8B Q5_K_M
# llama.cpp (Metal):  ~32 tok/s
# MLX:                ~38 tok/s
# Ollama (Metal):     ~30 tok/s (overhead from API layer)

The Real Trade-offs

Local models win on three axes: privacy, latency, and cost at scale. They lose on two: quality ceiling and convenience.

The quality gap between a local 7B model and GPT-4o or Claude 3.5 is real. On structured tasks — summarization, classification, extraction, code completion — a well-prompted 7B model at Q5 is close enough for production. On complex reasoning, creative writing, or tasks requiring broad world knowledge, the gap is significant.

The hybrid pattern wins: local for high-volume, low-stakes tasks; cloud for high-stakes, complex reasoning. Route based on task type, not ideology. You don't need GPT-4 to classify support tickets. You don't want a 7B model writing your legal briefs.

Setting Up a Production-Grade Local Stack

  1. Hardware: NVIDIA GPU with 12GB+ VRAM (RTX 3060 12GB is the budget entry point), or Apple M2 Pro+ with 32GB unified memory.
  2. Runtime: Ollama for simplicity, llama.cpp for control, vLLM for serving.
  3. Default model: Llama 3.1 8B Q5_K_M for general tasks, Qwen2.5-Coder 7B for code.
  4. Fallback: Cloud API (Claude, GPT-4) for tasks that exceed local capability.
  5. Routing: A simple classifier that sends tasks to local or cloud based on complexity signals.
# Example: Ollama setup with model pulling
ollama pull llama3.1:8b-q5_K_M
ollama pull qwen2.5-coder:7b-q5_K_M

# Serve with API compatibility
OLLAMA_HOST=0.0.0.0:11434 ollama serve

# Now hit it like OpenAI
curl http://localhost:11434/v1/chat/completions \
  -d '{"model":"llama3.1:8b-q5_K_M","messages":[{"role":"user","content":"Explain quantization"}]}'
"The question isn't whether local models are good enough. It's whether they're good enough for this specific task. Answer that per-task and you'll save a fortune."

The Bottom Line

Local models in 2026 are viable for the long tail of AI tasks — the high-volume, low-stakes work that would cost a fortune at cloud API rates. Set them up right (Q5_K_M, adequate VRAM, smart routing) and they handle 80% of the workload at 0% of the per-token cost. Route the remaining 20% to the best cloud model you can afford. That's the architecture that wins.

Z
Z.AI — GLM Models & Claude Code Support · partner
Access GLM-5, GLM-4, and 30+ models. Free tier available.
10% off →