by Matt Corbett
Chaos-engineering & resilience team — agents that make distributed systems survive failure. resilience-architect owns the DESIGN side: failure-mode analysis (FMEA), resilience patterns (timeouts, retries with backoff+jitter, bulkheads, circuit breakers, load shedding, graceful degradation, idempotency), the observability-maturity gate for whether you're ready to run chaos, and redundancy/DR (multi-AZ/region, failover, RTO/RPO). chaos-experiment-engineer owns the EXPERIMENT + VERIFICATION side: hypothesis-driven experiments, blast-radius containment (automatic abort/rollback), game days, fault injection (latency/error/resource/dependency-outage/partition/zone), and verifying the pattern held under load+fault. skills, a knowledge bank (decision trees + a dated 2026 patterns reference), templates. Distinct from observability-sre (metrics/SLO/on-call — a hard prerequisite), devops-cicd (progressive delivery), performance-engineering (load), incident-response-dfir. Requires ravenclaude-core@>=0.7.0.
Claude Code3 Skills