← Back to Blog

Context Window Strategies —

📅 🏷 Essay
Essay

Every AI conversation has a ceiling. The context window — the total tokens a model can process in a single session — is the invisible wall that kills long-running agent workflows. You hit it mid-task, the model forgets what it was doing, and you're back to square one. Here's how to manage it.

128K
tokens — GPT-4o context
200K
tokens — Claude 3.5
1M
tokens — Gemini 1.5
~60%
wasted on average per session

The Real Problem

Context windows are getting larger, but that's not the solution — it's the problem. Larger windows mean more rope to hang yourself with. Most sessions waste 60%+ of their context on:

The fix isn't bigger windows. It's smarter usage.

Strategy 1: Progressive Summarization

Don't keep the full history. After every N turns, compress the conversation into a summary that captures decisions, state, and open questions. The model gets a compressed past and a detailed present.

# Pseudocode for progressive summarization
if turn_count % COMPRESS_EVERY == 0:
    summary = model.summarize(history[-COMPRESS_EVERY:])
    context = [system_prompt, summary] + recent_turns
else:
    context = full_history

This trades accuracy for space. The summary loses detail, but keeps the critical path. In practice, summaries that include "what was decided" and "what's still open" retain 90% of what matters in 10% of the tokens.

Strategy 2: Selective Context Loading

Not every turn needs every file. Load context on demand based on what the model is actually doing:

ActionContext NeededTypical Size
Reading a fileJust that file2-5K tokens
Fixing a bugFile + related tests + error output5-15K tokens
RefactoringFile + all callers + type definitions15-50K tokens
Architecture changeMultiple files + docs + history50-100K+ tokens

Tools like LeanCTX implement this with "read modes" — signatures-only for browsing, full content for active editing. The difference is dramatic: a 50K-token codebase fits in 3K tokens when you only need function signatures.

Strategy 3: Sliding Window with Checkpoints

For agent workflows that run for 50+ turns, keep a sliding window of the last N turns plus periodic checkpoint summaries. This is how Claude Code and similar agents manage to run for hours without running out of context:

context = [
    system_prompt,
    checkpoint_summaries,  # Every 10 turns, compressed
    last_5_turns,         # Full detail for recent work
    current_task_state     # What we're doing right now
]
The key insight: the model doesn't need to remember exactly what it said 30 turns ago. It needs to remember what was decided and what's still pending.

Strategy 4: External Memory

When the conversation exceeds the window, offload to persistent storage. Write decisions to a file, store state in a database, use vector search for relevant context retrieval. This is how agents like Hermes operate across sessions — the context window is just the working memory, not the entire knowledge base.

The pattern:

  1. Before each turn: retrieve relevant context from external memory (vector search, keyword match)
  2. During the turn: work with retrieved context + recent history
  3. After the turn: store new decisions and findings back to external memory

Strategy 5: Token Budgeting

Treat your context window like a budget. Allocate tokens deliberately:

Never fill the context window past 80%. Models degrade in quality as they approach the limit. The last 20% of context produces noticeably worse output.
"The context window is not a storage container. It's a workspace. Keep it clear, load what you need, summarize what you don't."

The Bottom Line

Context management is the difference between an agent that works for 5 turns and one that works for 500. The five strategies — progressive summarization, selective loading, sliding windows, external memory, and token budgeting — can be combined. Start with budgeting, add selective loading, then progressively add the others as your workflows get more complex. The goal is always the same: maximum signal, minimum tokens.

Z
Z.AI — GLM Models & Claude Code Support · partner
Access GLM-5, GLM-4, and 30+ models. Free tier available.
10% off →