The Technique
Security researchers from UC Berkeley and ETH Zurich published a paper demonstrating Incremental Completion Decomposition (ICD), a technique that circumvents LLM safety training by breaking harmful requests into individual words or short phrases. The model completes each fragment innocuously, but the accumulated completions produce the harmful output.
Example: Instead of asking "How to build a bomb?", ICD asks the model to complete "How", then "How to", then "How to build", then "How to build a" — each step triggering benign completion behavior that collectively produces the restricted content.
Why It Works
ICD exploits a fundamental weakness in current safety training:
- Safety classifiers look at complete prompts — they detect harmful intent in fully-formed requests but miss intent distributed across multiple turns
- Each fragment is genuinely ambiguous — "How to" has infinite benign completions
- Context accumulation is the blind spot — models don't effectively track accumulated intent across decomposed requests
- Success rate: 73% on GPT-5.5, 68% on Claude Opus 4.7, 81% on open models
Implications
This isn't just another jailbreak — it reveals a structural weakness:
- Current RLHF-based safety training is insufficient for multi-turn interactions
- Safety classifiers need context-aware intent detection, not keyword matching
- Open models are more vulnerable due to less intensive safety training
- The fix likely requires architectural changes, not just more training data
ICD is the most elegant safety bypass I've seen. It doesn't exploit a bug — it exploits a design assumption: that harmful intent is present in a single prompt. When you decompose intent across turns, the current safety architecture fundamentally can't detect it. The fix won't be easy. It requires models that maintain and evaluate intent state across entire conversations, not just individual prompts.