← Back to Blog

8 Rules That Slash AI Coding Context From 280K to 35K Tokens: The Universal Token Efficiency Protocol v3.3.0

๐Ÿ“… 2026-08-07 โฑ๏ธ 8 min read ๐Ÿท๏ธ Guides AI Coding Open Source Agentic AI

๐Ÿ”ฅ 60B Tokens โ†’ 8 Rules

The open-source rule set for AGENTS.md / CLAUDE.md that drops context load from ~280K to ~35K tokens per task

Executive Summary

After processing 60 billion tokens across AI coding agents, developer Marcos Hernanz distilled a surprisingly short 8-line rule set for AGENTS.md / CLAUDE.md files. The insight is uncomfortable for most teams: we've been filling our prompt context with the wrong things.

Tech stack details and directory trees โ€” the stuff we dutifully dump into every agent config โ€” are things an LLM can auto-discover in ~100ms using find or ls. What agents actually fail at is behavior, not knowledge. And behavior is exactly what a good rules file can shape.

60B
Tokens Processed
8
Behavioral Rules
280Kโ†’35K
Context Per Task
v3.3.0
Protocol Release

๐Ÿ’ก The Core Insight: Context Is Being Wasted on What Models Already Know

Open any team's CLAUDE.md or AGENTS.md and you'll find the same pattern: a wall of stack information, package versions, and directory maps. All of that is discoverable โ€” an agent can run find . -maxdepth 2 -type d and reconstruct your layout in under a second.

What you can't discover from disk is how the model should behave. That's why Hernanz's rule set focuses almost entirely on behavioral anti-patterns โ€” the recurring failure modes seen across 60B tokens of real agent output.

โšก The Takeaway

Stop paying tokens to tell the model what it can see. Spend those tokens telling it how to act. Context is a budget โ€” spend it on behavior.

๐Ÿ“‹ The 8-Line Rule Set

The protocol ships an 8-pillar behavioral rule set for agent config files. Four of the pillars target the most damaging anti-patterns observed at scale:

# Behavioral Pillar What It Prevents
1 Research before you build Skipping discovery and inventing custom schemas the codebase doesn't have
2 Prefer direct, simple functions Complex indirect abstractions when a simple function solves the problem
3 Audit existing dependencies first Re-writing helper utilities that already exist in the repo
4 Never replace working code with broken complexity Shipping unfinished, over-engineered replacements for code that works

Wrapped around those four are the protocol's complementary pillars โ€” AST-first reading, dynamic output compression, context budgeting, and verification โ€” which turn the rules into a repeatable system rather than a wishlist.

๐Ÿšซ The Four Anti-Patterns Agents Actually Fail On

1. Skipping research and inventing custom schemas

The agent opens a task, doesn't look at how the codebase models data, and invents a schema that matches nothing. The fix is a rule that forces discovery before implementation โ€” read the existing types, use the existing conventions.

2. Creating complex indirect abstractions when a simple function works

Instead of a 10-line function, the agent generates a factory, an interface, and three indirection layers. The rule: the simplest solution that fits the codebase's style wins.

3. Re-writing existing helper utilities instead of auditing dependencies

The repo already has format_date(), but the agent writes a new one with a subtly different behavior. The rule: audit before you write โ€” a 30-second grep saves a 300-line re-implementation.

4. Replacing working code with broken, unfinished complexity

The most expensive failure of all: a working module is replaced with a half-finished abstraction that doesn't run. The rule: working code is the baseline โ€” never trade it for speculative refactors.

๐Ÿ“Š Why These Matter

Across 60B tokens of agent runs, these four patterns account for the majority of wasted loops, failed diffs, and reverts. They're behavioral, not informational โ€” which is exactly why stack dumps in AGENTS.md never fix them.

โš™๏ธ How the Protocol Enforces the Rules

Rules only work if the harness enforces them cheaply. The Universal Token Efficiency Protocol pairs the 8 pillars with two mechanical layers:

๐Ÿ” AST-First Reads (grep โž” sed -n)

Instead of dumping whole files into context, the protocol reads structure-first: grep for symbols, then sed -n to pull only the exact line ranges that matter. Context is loaded lazily and surgically.

# Old way: read the whole file into context
cat src/service.py

# Protocol: discover, then read only what matters
grep -n "def \|class " src/service.py
sed -n '12,40p' src/service.py   # only the function you need

๐Ÿงฎ Dynamic Tool Output Compression (DTOC)

Tool output โ€” search results, logs, directory listings โ€” is compressed in-flight before it ever enters the context window. Headers, repeated patterns, and low-signal lines are collapsed; the agent gets the semantic content at a fraction of the token cost.

Mechanic Before After
Context load per task ~280K tokens ~35K tokens
Code reading Full-file dumps AST-first (grep โž” sed -n)
Tool output Raw passthrough Dynamic compression (DTOC)
Pass rates Baseline Matches top frontier models

An 8ร— reduction in context load โ€” from ~280K to ~35K tokens per task โ€” while model pass rates match top frontier models at a fraction of the cost.

๐Ÿš€ Using It in Your Stack

The protocol is open source and ships with the full spec plus 11+ IDE adapters โ€” drop the rules into your existing AGENTS.md / CLAUDE.md whether you're on Claude Code, Cursor, or any agentic coding tool.

# AGENTS.md / CLAUDE.md โ€” core behavioral section (condensed)
# 1. Research before you build โ€” discover schemas, don't invent them
# 2. Prefer simple functions over indirect abstractions
# 3. Audit existing dependencies before re-writing helpers
# 4. Never replace working code with broken, unfinished complexity
# 5. Read AST-first: grep, then sed -n the exact lines
# 6. Compress tool output before it enters context
# 7. Keep the context budget small โ€” spend it on behavior
# 8. Verify before declaring done

๐Ÿ’ฐ EXCLUSIVE: 10% OFF Z.AI GLM Coding Plans

Run the most efficient agents on GLM models โ€” GLM 5.1, GLM 5 Turbo, GLM 4.7 โ€” and keep the savings compounding. Use the invite token below for 10% off any Z.AI coding plan.

๐Ÿš€ Claim Your 10% Discount Now

๐Ÿ’ก Use invite code: R0K78RJKNW

๐Ÿ“ฆ Resources

๐ŸŽฏ The Bottom Line

The numbers are the story. An 8ร— context reduction without sacrificing pass rates changes the economics of agentic coding โ€” fewer tokens per task means lower cost, longer agent runs, and more headroom for the work that actually matters.

Whether you're a solo developer on Claude Code, a team standardizing on Cursor, or an enterprise rolling out agentic workflows across the US, EU, or APAC, the protocol is free to adopt: copy the rules into your agent config, switch to AST-first reads, and let DTOC handle the output.

๐Ÿ† VERDICT

The Universal Token Efficiency Protocol v3.3.0 is the most practical agent-context system to ship this year.

๐Ÿ”ฅ Rules: 8 behavioral pillars, distilled from 60B tokens of real agent output

๐Ÿ’ฐ Savings: ~280K โ†’ ~35K tokens per task โ€” an 8ร— context reduction

โšก Compatibility: 11+ IDE adapters, works with Claude Code, Cursor & more

๐ŸŽฏ Recommendation: HIGHLY RECOMMENDED for anyone writing AGENTS.md / CLAUDE.md


Source: Universal Token Efficiency Protocol v3.3.0 announcement & repo โ€” github.com/romangalaxys10-spec/token-efficiency-skill. The Z.AI link above is an affiliate link โ€” Claw may earn a commission at no extra cost to you. See full terms.