by Matt Corbett
AI/LLM red-teaming team — agents (ai-redteam-lead, adversarial-testing-engineer) for the layer answering 'can this AI system be made to do harm, leak data, or exceed its authority — and how do we harden it?': threat modeling + rules of engagement, the attack taxonomy (OWASP LLM Top 10 2026 + MITRE ATLAS), direct vs indirect prompt injection, jailbreaks (roleplay/encoding/many-shot/crescendo), data exfiltration & training-data extraction, agentic tool-abuse / excessive agency, multimodal attacks, and defense-in-depth remediation. Fluent in automated red-team harnesses (PyRIT, Garak, Promptfoo red-team, Giskard) and likelihood×impact severity. skills, a knowledge bank (attack-taxonomy decision tree + 2026 patterns), and templates. Distinct from llm-evaluation-engineering (quality-regression eval), trust-and-safety (content-moderation / T&S policy), and security-engineering (app/infra pentest) — the adversarial AI-security layer over model- and agent-based systems. Requires ravenclaude-core@>=0.7.0.
Claude Code3 Skills