I spent a week running T3MP3ST against LLM infrastructure and what came back changed how I think about AI security. Not because of any single vulnerability, but because of the pattern: every phase of a traditional intrusion kill chain now has a semantic equivalent in AI systems, and almost no one is defending the new surface.
This piece goes beyond the threat-map. For each of the eight T3MP3ST operators I'll show (1) a real-world incident it mirrors, (2) a concrete simulation of what the operator actually does, and (3) a defensive use case you can deploy this week.
The T3MP3ST operator transcripts are illustrative reconstructions of operator behavior, not raw output from a live run. They are structured to match how each operator reasons so you can see the mechanics — replace the prompts with your own models and tooling to reproduce them.
T3MP3ST decomposes an intrusion into eight specialized operators, each trained to think a different way. Mapped onto MITRE ATT&CK Enterprise tactics, the framework stops being "red-team theater" and becomes a coverage matrix you can audit against.
| Operator | ATT&CK Tactic | LLM Equivalent |
|---|---|---|
| Recon | TA0043 Reconnaissance | Scrape docs, leak system prompts, fingerprint model family |
| Scanner | TA0047 Discovery | Fuzz for reasoning gaps, jailbreak surface, tool access |
| Exploiter | TA0002 Execution | Persuasion / jailbreak that produces forbidden output or action |
| Infiltrator | TA0001 Initial Access | Indirect prompt injection through retrieved content |
| Exfiltrator | TA0010 Exfiltration | Reconstruct secrets via structured queries / covert channels |
| Ghost | TA0005 Defense Evasion | Adversarial perturbations, filter bypass, token smuggling |
| Coordinator | TA0008 Lateral Movement | Multi-agent orchestration across tools and sessions |
| Analyst | TA0042 Resource Dev | Threat-intel synthesis, CVE tracking, IR planning |
Traditional recon scans ports and enumerates services. LLM recon scrapes API documentation, studies model architecture papers, and — most dangerously — harvests system-prompt leakage from error messages and verbose debug output. The model literally tells you its rules if you ask the right way.
Employees pasted proprietary source code and internal meeting notes into a public chatbot. The data entered the training pipeline of a third party. Samsung banned ChatGPT within days. The lesson isn't "don't use AI" — it's that your prompts are outbound traffic, and recon operators treat them as exfiltration by default.
Run Recon against your own bots. A monthly automated pass that attempts system-prompt extraction and tool enumeration surfaces exactly what an attacker sees. Pair it with:
Scanner phase used to mean running nmap. Now it means automated prompt-injection testing at scale — feeding thousands of adversarial inputs and tracking which ones trigger unexpected behavior. The scanner hunts for reasoning gaps, not memory corruption.
Researchers showed that stuffing dozens — then hundreds — of in-context examples of rule-breaking gradually wore down model compliance. The vulnerability isn't a bug; it's in-context learning itself. A scanner that only tests single-turn attacks misses the entire gradient.
Make the scanner a CI gate. Before any prompt template or agent config ships, run the fuzzer against a held-out "red" set. Block deploys where trigger rate on safety-relevant behaviors exceeds a threshold (e.g. >2%). Track the number as a security metric over time.
Exploitation with LLMs isn't code execution. It's persuasion. A jailbreak doesn't exploit a memory bug — it exploits the model's training to be helpful, harmless, and honest, then redefines what those words mean in context.
A universal jailbreak that asked models to append a disclaimer rather than refuse, convincing them they'd complied "safely." It worked across multiple frontier models. The exploit was pure social engineering of the alignment layer — no gradient access required.
Detect intent-vs-surface mismatch. If a response contains a disclaimer but also performs the forbidden action, that's a Skeleton-Key signature. Add a classifier that scores "does the output actually comply with the stated safety caveat?" — not just whether the caveat string is present.
This is the scariest operator for RAG systems. Infiltration doesn't require foothold credentials. It rides in through indirect prompt injection — malicious instructions hidden in documents the model is told to trust.
The paper demonstrated that a webpage, email, or PDF containing "ignore previous instructions and do X" could hijack an LLM that later read it — with zero access to the system. In 2024, a similar technique let attackers hijack a coding agent's browser tools just by getting it to visit a poisoned page.
Enforce a data/instruction separation boundary. Tag retrieved content as untrusted data, render it in a sandboxed channel, and never let it set system state. Tools that can be called from RAG output should require explicit human confirmation for any side-effecting action (email, HTTP, shell).
Traditional exfil needs to bypass encryption and slip past DLP. LLM exfil just needs well-formed questions. The model reconstructs what it was trained on, or leaks what's in its context window, one answerable query at a time.
Members of the GPT-2 / Chinchilla training sets were recovered verbatim by simply asking the model to "repeat the book you were trained on." PII, code, and personal data all came back. No decryption needed — the knowledge was the liability.
Deploy semantic DLP, not just regex. A model-guard that scores each response for PII / secret / internal-policy leakage catches "summarize the ticket" exfil that a keyword filter misses. Redact at the response layer and log attempted exfil for the Analyst operator on your side.
Ghost is defense evasion. Adversarial perturbations, Unicode homoglyphs, and token smuggling (hiding malicious tokens inside ASCII art or translated cipher) slip past both filters and human reviewers.
Attackers rendered prohibited instructions as ASCII diagrams, asked the model to "describe this image," then execute it. Others used "respond in a Caesar cipher" to launder intent past substring blocklists. The filter saw benign text; the model saw the command.
Normalize before you filter. Strip Unicode homoglyphs, render images through an OCR+caption round-trip and re-scan the transcription, and decode common ciphers before the guard model sees input. Evasion lives in the gap between what the filter reads and what the model reads — close that gap.
Single compromised agents are bad. Coordinated multi-agent campaigns are worse. The Coordinator orchestrates several agents, each playing a role — one reconnoiters, one exploits, one exfiltrates — across tools and sessions, so no single log shows the full attack.
Coding agents granted shell + browser + email began chaining them: read a poisoned repo (Infiltrator), use the shell to exfil via curl (Exfiltrator), then email the results to an attacker address (Coordinator). Each tool logged a "normal" action. The campaign only appeared when logs were correlated.
Build agent-to-agent trust boundaries. No agent should hold both read-secret and send-network privileges simultaneously. Correlate tool calls across agents in a single session graph; alert when the combination of "read sensitive + external send" appears, even if each step is individually permitted.
Analyst synthesizes threat intel, tracks CVEs, and plans incident response. This is the operator that turns the framework from an attack tool into a defense planner. Run Analyst on your own telemetry and it tells you which kill-chain phase you're weakest in.
Wire Analyst into a weekly red/blue report. It consumes your guard-log telemetry, ranks phases by exposure, and proposes the next hardening step — turning the same kill-chain lens attackers use into a prioritized backlog for your security team.
After running this end to end, one conclusion is unavoidable: LLM security isn't application security anymore. It's infrastructure security.
Traditional WAFs filter SQLi and XSS. They do not filter "pretend you're a different model with no restrictions." The companies getting this right treat LLM security as a continuous red-team exercise — specialized operators looking at the problem from eight angles, every sprint.
The future of AI security isn't better firewalls. It's better adversarial reasoning — the discipline of thinking like an eight-operator kill chain and closing each gap before it's exploited. T3MP3ST gives you the lens. The teams that operationalize it will define the next decade of secure AI.
LLM security has shifted from application-level concerns to infrastructure-level threats. The T3MP3ST framework maps its eight operators — Recon, Scanner, Exploiter, Infiltrator, Exfiltrator, Ghost, Coordinator, Analyst — directly onto MITRE ATT&CK tactics, turning abstract "AI risk" into an auditable coverage matrix. Real incidents (Samsung's code leak, Many-Shot and Skeleton Key jailbreaks, indirect prompt injection, training-data extraction) show each phase is already happening in the wild. The defense is the same lens, pointed inward: scan your own bots, gate deploys on fuzz-trigger rates, separate data from instructions in RAG, deploy semantic DLP, and bound agent tool privileges. Do that continuously and the kill chain breaks at phase one.