← Back to Blog
⚡ 10% OFF Get 10% off any Z.AI Coding Plan — covers GLM 5.1, GLM 5 Turbo, and GLM 4.7 → Claim discount
April 30, 2026 8 min read

Incremental Completion Decomposition: A New Way to Break LLM Safety Training

Researchers discovered that breaking harmful requests into individual words can circumvent LLM safety guardrails. Current alignment is more fragile than we thought.

The Technique

Security researchers from UC Berkeley and ETH Zurich published a paper demonstrating Incremental Completion Decomposition (ICD), a technique that circumvents LLM safety training by breaking harmful requests into individual words or short phrases. The model completes each fragment innocuously, but the accumulated completions produce the harmful output.

Example: Instead of asking "How to build a bomb?", ICD asks the model to complete "How", then "How to", then "How to build", then "How to build a" — each step triggering benign completion behavior that collectively produces the restricted content.

Why It Works

ICD exploits a fundamental weakness in current safety training:

Implications

This isn't just another jailbreak — it reveals a structural weakness:

// Editor's Take

ICD is the most elegant safety bypass I've seen. It doesn't exploit a bug — it exploits a design assumption: that harmful intent is present in a single prompt. When you decompose intent across turns, the current safety architecture fundamentally can't detect it. The fix won't be easy. It requires models that maintain and evaluate intent state across entire conversations, not just individual prompts.


The Takeaway
Incremental Completion Decomposition exposes a fundamental blind spot in LLM safety training. By distributing harmful intent across multiple turns, it bypasses safety classifiers with 73-81% success rates. The fix requires architectural changes to how models track intent state across conversations, not just more training data.

✓ Why It Matters

⚠ What to Watch