#agent-security
-
GPT Security: Attack Surfaces and Production Controls
A technical guide to GPT security covering prompt injection, custom GPTs, agent actions, data handling, and layered production controls.
-
ChatGPT Exploits: Injection, Memory Hijack, Plugin Abuse
How ChatGPT gets exploited: indirect prompt injection, CSRF memory poisoning, cross-plugin request forgery, agent tool abuse, and the controls that help.
-
Insecure Output Handling: LLM05:2025 Attacks and Defenses
Insecure output handling turns model text into XSS, SQL injection, or remote code execution. The attack chain, the CVEs it produced, and the controls.
-
LLM Security Vulnerabilities: What Actually Gets Exploited
The LLM security vulnerabilities showing up in real CVEs: prompt injection, system prompt leakage, RAG poisoning, and tool-call bugs that turn into RCE.
-
GPT-Red: What OpenAI's Prompt-Injection Attacker Proves
OpenAI says GPT-Red beat human red teamers 84% to 13% on prompt injection. Here is what those numbers measure, and what they leave unanswered.
-
LLM Attack Taxonomy: Prompt Injection, Jailbreaks, Agent Hijack
A practitioner's map of LLM attack classes: direct and indirect prompt injection, jailbreaks, RAG poisoning, and agent tool-call abuse, mapped to OWASP.
-
AI Red Team: Methodology, Tooling, and Attack Surface
A practitioner's guide to AI red teaming: how LLM attack surface differs from traditional app testing, and the techniques and tooling that map it.
-
Prompt Injection in 2025: OpenAI vs. Broken Defenses
OpenAI's advisory on prompt injection landed the same week research showed adaptive attacks beat published defenses more than ninety percent of the time.
-
The Audit Gap: Why Red-Teaming Can't Certify Governance Claims
A position paper formalizes the mismatch between what AI governance frameworks ask evaluators to verify and what behavioral red teaming can actually show.
-
Prompt Injection Examples: A Practitioner's Attack Library
A technical breakdown of real prompt injection examples across direct, indirect, multimodal, and RAG-poisoning attacks, with payloads and conditions.
-
LLM Prompt Injection: Taxonomy, Real Patterns, and Defenses
A technical breakdown of LLM prompt injection: direct, indirect, and agent-targeting variants, the attack patterns seen in the wild, and defenses that work.
-
Prompt Hacking: Taxonomy, Techniques, and What Works on LLMs
A practitioner breakdown of prompt hacking: the three attack families of injection, leaking, and jailbreaking, how each works, and what defenses hold.
-
Prompt Injection Attack: Techniques, Variants, and Defenses
A practitioner's breakdown of prompt injection attacks: direct, indirect, and multi-modal, covering the HouYi framework, real CVEs, and mitigations that hold.
-
LLM Security: A Practitioner's Map of the Attack Surface
What LLM security means in 2026: the attack classes red teamers test, the controls that hold up under fire, and the frameworks that map the territory.