#adversarial-ml
-
LLM Attack Taxonomy: Prompt Injection, Jailbreaks, Agent Hijack
A practitioner's map of LLM attack classes: direct and indirect prompt injection, jailbreaks, RAG poisoning, and agent tool-call abuse, mapped to OWASP.
-
The Adversarial ML Attack Taxonomy: A Red Teamer's Reference
A working taxonomy of attacks against ML systems, covering evasion, poisoning, privacy, and abuse, mapped to attacker access and aligned to NIST and ATLAS.
-
The Audit Gap: Why Red-Teaming Can't Certify Governance Claims
A position paper formalizes the mismatch between what AI governance frameworks ask evaluators to verify and what behavioral red teaming can actually show.
-
Prompt Injection Examples: A Practitioner's Attack Library
A technical breakdown of real prompt injection examples across direct, indirect, multimodal, and RAG-poisoning attacks, with payloads and conditions.
-
AI Red Teaming Hub: Your Guide to Offensive AI Security
The central index for offensive AI security on this site: prompt injection, jailbreaks, adversarial ML, red team methodology, and the tooling behind it.
-
Automated Jailbreak Attacks and the Transfer Problem
How automated attack generation works — PAIR, GCG, and TAP — why jailbreaks port across model families, and what that does to an assessment's threat model.
-
LLM Bypass: How Attackers Circumvent Safety Alignment by Layer
A technical breakdown of LLM bypass techniques: adversarial suffixes, shallow alignment exploits, fine-tuning attacks, and guardrail evasion, layer by layer.
-
LLM Jailbreak: Attack Taxonomy, Techniques, and Defense Reality
A technical breakdown of LLM jailbreak attack classes: many-shot, Crescendo multi-turn escalation, roleplay, and encoding, plus what defense really achieves.
-
Prompt Hacking: Taxonomy, Techniques, and What Works on LLMs
A practitioner breakdown of prompt hacking: the three attack families of injection, leaking, and jailbreaking, how each works, and what defenses hold.
-
Prompt Injection Attack: Techniques, Variants, and Defenses
A practitioner's breakdown of prompt injection attacks: direct, indirect, and multi-modal, covering the HouYi framework, real CVEs, and mitigations that hold.
-
ChatGPT Jailbreak Prompt Taxonomy: Classes, Rates, and Defenses
A research-grounded breakdown of ChatGPT jailbreak prompt classes, from DAN and persona injection to multi-turn escalation, and what the defenses catch.
-
OSCP and CEH in 2026: What Carries Over to AI Red Teaming
A free OSCP and CEH study offer raises the question every pentester should answer: which of those skills transfer to AI red teaming, and which do not.