Every AI conversation has a ceiling. The context window — the total tokens a model can process in a single session — is the invisible wall that kills long-running agent workflows. You hit it mid-task, the model forgets what it was doing, and you're back to square one. Here's how to manage it.
Context windows are getting larger, but that's not the solution — it's the problem. Larger windows mean more rope to hang yourself with. Most sessions waste 60%+ of their context on:
The fix isn't bigger windows. It's smarter usage.
Don't keep the full history. After every N turns, compress the conversation into a summary that captures decisions, state, and open questions. The model gets a compressed past and a detailed present.
# Pseudocode for progressive summarization
if turn_count % COMPRESS_EVERY == 0:
summary = model.summarize(history[-COMPRESS_EVERY:])
context = [system_prompt, summary] + recent_turns
else:
context = full_history
This trades accuracy for space. The summary loses detail, but keeps the critical path. In practice, summaries that include "what was decided" and "what's still open" retain 90% of what matters in 10% of the tokens.
Not every turn needs every file. Load context on demand based on what the model is actually doing:
| Action | Context Needed | Typical Size |
|---|---|---|
| Reading a file | Just that file | 2-5K tokens |
| Fixing a bug | File + related tests + error output | 5-15K tokens |
| Refactoring | File + all callers + type definitions | 15-50K tokens |
| Architecture change | Multiple files + docs + history | 50-100K+ tokens |
Tools like LeanCTX implement this with "read modes" — signatures-only for browsing, full content for active editing. The difference is dramatic: a 50K-token codebase fits in 3K tokens when you only need function signatures.
For agent workflows that run for 50+ turns, keep a sliding window of the last N turns plus periodic checkpoint summaries. This is how Claude Code and similar agents manage to run for hours without running out of context:
context = [
system_prompt,
checkpoint_summaries, # Every 10 turns, compressed
last_5_turns, # Full detail for recent work
current_task_state # What we're doing right now
]
When the conversation exceeds the window, offload to persistent storage. Write decisions to a file, store state in a database, use vector search for relevant context retrieval. This is how agents like Hermes operate across sessions — the context window is just the working memory, not the entire knowledge base.
The pattern:
Treat your context window like a budget. Allocate tokens deliberately:
"The context window is not a storage container. It's a workspace. Keep it clear, load what you need, summarize what you don't."
Context management is the difference between an agent that works for 5 turns and one that works for 500. The five strategies — progressive summarization, selective loading, sliding windows, external memory, and token budgeting — can be combined. Start with budgeting, add selective loading, then progressively add the others as your workflows get more complex. The goal is always the same: maximum signal, minimum tokens.