Agent Plugins Marketplace
← All plugins

save-toolkit

v0.50.0

Application-engineering and site-reliability agents and reusable skills.

Claude CodeAgent Plugins29 Skills

By latent-sreLicense: MIT0 GitHub starsUpdated 1 hour ago

Directory evidence

Runtimes
Claude Code and Agent Plugins
Parsed components
29 skill or MCP entries
Source updated
Oct 1, 2026
Manifest status
Canonical path parsed

The directory validates manifest shape and source location. It does not execute the plugin or provide a security endorsement. Review the indexing methodology →

Install save-toolkit for Claude Code

Installs for the current user
claude plugin marketplace add IchenDEV/agent-plugin-mkt
claude plugin marketplace update agent-plugin-marketplace
claude plugin install save-toolkit@agent-plugin-marketplace

Paste and run these commands in a terminal with Claude Code. They add and refresh the PluginsMP catalog, then install this plugin.

The installer fetches third-party code from the source repository shown on this page. This directory validates manifest structure and source location, but does not perform a security audit; review the manifest, components, and source before installing.

Get the source manually
git clone https://github.com/latent-sre/save-toolkit

Clone the source repository, then follow its setup instructions to add the plugin to a compatible client. The repository root is the plugin root.

Plugin files

save-toolkit/
├── .claude-plugin/plugin.json
├── plugin.json
├── skills/agent-authoring/SKILL.md
├── skills/akamai-edge/SKILL.md
├── skills/backend-craft/SKILL.md
├── skills/ci-actions/SKILL.md
├── skills/database-reliability/SKILL.md
├── skills/eng-ladder/SKILL.md
├── skills/frontend-craft/SKILL.md
├── skills/gcp-ops/SKILL.md
├── skills/grafana/SKILL.md
├── skills/incident-investigation/SKILL.md
├── skills/obs-alerting/SKILL.md
├── skills/obs-dashboards/SKILL.md
├── skills/obs-logs/SKILL.md
├── skills/obs-metrics/SKILL.md
├── skills/obs-pipeline/SKILL.md
├── skills/obs-traces/SKILL.md
├── skills/operational-learning/SKILL.md
├── skills/operator-cli/SKILL.md
├── skills/pcf-deploy/SKILL.md
├── skills/pcf-ops/SKILL.md
├── skills/postmortem/SKILL.md
├── skills/production-change-gate/SKILL.md
├── skills/python-craft/SKILL.md
├── skills/resilience-analysis/SKILL.md
├── skills/root-cause/SKILL.md
├── skills/runbook/SKILL.md
├── skills/service-lifecycle/SKILL.md
├── skills/stack-profile/SKILL.md
└── skills/toil-reduction/SKILL.md

Included Skills29

agent-authoringskills/agent-authoring/SKILL.md

Create, repair, or security-review LLM-facing prompts, agents, skills, tool descriptions, graders, bounded Loop Engineering for evaluation/verification, and agent roster/delegation graphs. Triggers: 'write me an agent/skill/prompt', 'my skill fires too often', 'the output is the wrong shape', 'is this agent safe / prompt injection'. Not for source-code dependency, knowledge, or GraphRAG graphs, implementing a graph runtime, or the design contract for an executable workflow/state graph (agent-engineer's graph tier).

akamai-edgeskills/akamai-edge/SKILL.md

Akamai edge work in three lanes — triage (edge vs origin, Reference # error strings, cache status, WAF denials, DataStream 2), delivery config (Property Manager versions, staging-first activation, fast fallback), and mPulse RUM (network-side vs app-side slowdowns). Triggers: 'is it the CDN or the origin', 'an Akamai error page with Reference #<n>.<hex>.<epoch>.<hex>', 'is the WAF blocking real users', 'mPulse shows slow pages'. Not for backend log queries (obs-logs) or a firing alert (incident-investigation).

backend-craftskills/backend-craft/SKILL.md

Builds or changes an API or backend service — HTTP endpoints, workers, schedulers, the service behind a UI — and consumes third-party APIs safely (clients, SDK wrappers, sync jobs, webhooks), including our platform/obs APIs. Triggers: 'add an endpoint', 'wrap X behind an API', 'write a client for Y'. Not for UI work (frontend-craft), live-data operations (database-reliability).

ci-actionsskills/ci-actions/SKILL.md

Review, design, troubleshoot, and optimize GitHub Actions workflows: fast feedback, caching, test matrices, reusable jobs, reproducible artifacts, and secure delivery. Includes advice-only workflow reviews when the caller asks for findings without edits. Triggers: 'set up CI', 'speed up this pipeline', 'why is this workflow failing', 'harden the pipeline'. Not for application failures after deployment (pcf-ops, gcp-ops), runtime bugs (root-cause), or a live outage, including one a deploy caused (incident-investigation).

database-reliabilityskills/database-reliability/SKILL.md

Diagnose and improve data-layer reliability: slow queries, lock contention, replication lag, connection pools, schema migrations, and recovery evidence. Triggers: 'this query is slow', 'plan this schema migration', 'the connection pool is exhausted'. Not for app-side triage (pcf-ops), burn alerts (obs-alerting), persistence code (backend-craft).

eng-ladderskills/eng-ladder/SKILL.md

Select the engineering altitude for implementation, design, review, or growth feedback when work may span components, teams, migrations, or hard-to-reverse choices. Triggers: 'how rigorous should this be', 'review this at the principal level', 'is this a design doc or just a PR'. A scoped change with an obvious owner and existing pattern routes straight to its builder or craft skill. Active-alert troubleshooting belongs to incident-investigation.

frontend-craftskills/frontend-craft/SKILL.md

Build or change a web UI — pages, dashboards-as-app-features, forms, admin panels — from a single page to a full SPA, including serving it on PCF. Owns UI-layer TypeScript/React idiom — component state, interaction, accessibility, resilience UX. Triggers: 'build a UI for', 'add a page/form/table', 'make this dashboard page'. Not for the service behind the UI (backend-craft) or Grafana/observability dashboards (obs-dashboards to design, grafana to operate).

gcp-opsskills/gcp-ops/SKILL.md

Investigate application-side GCP failures during the migration — Cloud Run services and revisions, gcloud logging reads, what-changed correlation against revision deploys, and the project-vs-platform boundary. Triggers: 'the Cloud Run service is 503ing', 'read the GCP logs', 'container failed to listen on PORT', 'roll back the Cloud Run service to the previous revision'. Related owners: `stack-profile` (boundary and runtime decisions), obs-logs (log-query dialects), pcf-ops (the PCF side while both runtimes coexist).

grafanaskills/grafana/SKILL.md

Operate Grafana: find, explain, create, and edit dashboards; inspect alert state and notification paths; create or update Grafana-managed alert rules, and manage temporary silences. Triggers: 'explain this Grafana dashboard', 'create a Grafana alert', 'silence this alert', 'edit this Grafana dashboard'. Dashboard design uses obs-dashboards; alert/SLO design uses obs-alerting. Active-incident diagnosis belongs to incident-investigation.

incident-investigationskills/incident-investigation/SKILL.md

Helps a human SRE investigate a live incident, understand evidence, and choose the next useful step. Use for new pages, ongoing troubleshooting, interpreting supplied logs, graphs, metrics, traces or alerts, comparing mitigation options and recommending what to do, checking recovery, and preparing investigation handovers for an existing bridge/TLC (Techline Chat). Also explains operational signals outside a live incident. Supports first responders who do not know where to start, through experienced SREs. Triggers: 'I just got paged, what do I do', 'customers are reporting errors, where do I start', 'walk me through this incident', 'what should I check next'. Not for a delegated read-only lookup or investigation (sre-assistant agent), running incident command, stakeholder communications, or authoring postmortems (scribe).

obs-alertingskills/obs-alerting/SKILL.md

Design alerting that pages on symptoms — SLIs/SLOs and multi-window burn rates, Splunk saved-search alerts, Moogsoft correlation, and ThousandEyes synthetics. Triggers: 'define an SLO', 'this alert is too noisy', 'what should page', 'design a synthetic check'. Not for queries (obs-metrics, obs-logs) or dashboards (obs-dashboards).

obs-dashboardsskills/obs-dashboards/SKILL.md

Design dashboards around the on-call reader's questions: service health, golden signals, useful panels, units, comparisons, missing-data presentation, and drill-downs. Triggers: 'design a dashboard', 'what should we dashboard', 'which panels do we need', 'make this dashboard easier to read'. Grafana reads, saves, JSON/API details, and alert operations belong to grafana; alert/SLO design belongs to obs-alerting. Not for dashboards built into an application UI (frontend-craft).

obs-logsskills/obs-logs/SKILL.md

The answer is in the logs — find error spikes, read them over time, correlate one request across services, compare before/after a deploy. Backends: Splunk (SPL), Loki (LogQL), and Cloud Logging on GCP — the reference teaches the dialect. Triggers: 'search the logs', 'why are there 500s', 'write a log query', 'follow this correlation id'. Ownership map only: obs-metrics owns metrics, obs-dashboards owns dashboard design, grafana owns Grafana operations, and obs-alerting owns alert design. Deciding what a live page means or what to do next belongs to incident-investigation, which routes query work here.

obs-metricsskills/obs-metrics/SKILL.md

The answer is in the metrics — latency percentiles, error ratios, saturation, rates, missing-data traps. Backends: Wavefront (WQL), Mimir/Prometheus (PromQL), and Cloud Monitoring on GCP (PromQL; MQL is deprecated). Triggers: 'query the metrics', 'graph the error rate', 'is latency up', 'write a metric alert query'. Not for alert design (obs-alerting) or logs (obs-logs). Deciding what a live page means or what to do next belongs to incident-investigation, which routes query work here.

obs-pipelineskills/obs-pipeline/SKILL.md

What ships telemetry where — instrument a service with OTel and route metrics, traces, and structured logs through Alloy/collectors to Loki, Mimir, and Tempo. Triggers: 'instrument this service', 'add telemetry', 'logs are not reaching Loki', 'configure the Alloy exporter'. Not for reading signals (obs-logs, obs-metrics, obs-traces) or Grafana datasource/dashboard configuration (grafana).

obs-tracesskills/obs-traces/SKILL.md

Follow one request across services — when logs say 'slow' and metrics say 'sometimes', the trace says where. Read waterfalls, find the span that ate latency, and correlate trace ids with logs. Backends: Tempo (TraceQL) and Cloud Trace on GCP. Triggers: 'trace this request', 'where did the latency go', 'open this trace id'. A request or correlation id with no trace id starts in obs-logs. Not for trace instrumentation (obs-pipeline).

operational-learningskills/operational-learning/SKILL.md

Turn a resolved incident, drill, audit, or approved component or alert change into durable knowledge (service and alert cards, runbook dispositions), and answer ownership questions from those records. Triggers: 'which team owns payments and how do I page them', 'what depends on ledger', 'capture durable operational lessons', 'apply the operational-learning closeout'. Direct KB writing belongs to scribe, which selects closeout mode and applies this skill; active incidents route to incident-investigation, alert design to observability-engineer, and fleet prompt failures to agent-engineer. Retrospective write-ups use postmortem.

operator-cliskills/operator-cli/SKILL.md

Build or repair an operator-facing CLI that works interactively and in automation, including partial failures, machine output, dry runs, and interruption. Triggers: 'build an operator CLI', 'make this command safe for automation', 'fix this CLI contract'. Not for running an existing operational command (production-change-gate; during a live incident, incident-investigation), a web UI (frontend-craft), or the service behind a CLI (backend-craft).

pcf-deployskills/pcf-deploy/SKILL.md

Plan human-approved VMware TAS/PCF application deploys, blue-green cutovers, scaling, and rollback verification. Triggers: 'deploy this app to PCF', 'design a blue-green deploy', 'scale this PCF app'. Not for readiness (production-change-gate) or incident mitigation advice (incident-investigation).

pcf-opsskills/pcf-ops/SKILL.md

Investigate application-side PCF/TAS failures with cf app, events, logs, and routes, and distinguish app faults from platform-wide symptoms. Triggers: 'the app is crashing', 'why is my app 502-ing', 'exit code 137', 'X-Cf-RouterError'. Not for widespread Diego/Gorouter failures, which go to the platform team with evidence.

postmortemskills/postmortem/SKILL.md

Write a blameless postmortem for a resolved incident: systemic causes, timeline, detection, response, and owned action items in the standard structure. Triggers: 'write a postmortem for INC-1234', 'draft the retro for the order-router outage', 'apply the postmortem structure'. Direct retrospective writing belongs to scribe, which selects postmortem mode and applies this skill; active incidents route to incident-investigation; ITO runs coordination in the TLC.

production-change-gateskills/production-change-gate/SKILL.md

Gate the path from code to production with one checklist per question: is this change ready to merge, is this build ready to ship, and may this exact action run against production. Records a human decision and never executes. Triggers: 'is this ready to merge', 'is this build ready to ship', 'can I run this cf command in prod', 'authorize this production change'. Not for code review itself (reviewer) or for incident mitigation advice (incident-investigation).

python-craftskills/python-craft/SKILL.md

Write, refactor, or modernize Python in services, CLIs, scripts, libraries, and tests, routine or difficult: simplify structure, reduce duplicated logic, fix established defects, and replace custom machinery with suitable maintained libraries. Load before editing any Python; also use for explanations and practical starting designs. Triggers: 'improve this Python', 'refactor this Python', 'modernize this module', 'explain this Python'. Not for operating a running service or changing another language.

resilience-analysisskills/resilience-analysis/SKILL.md

Analyze failure propagation and design service resilience using source, configuration, operational evidence, and explicit assumptions. Triggers: 'what happens if this dependency slows down', 'find resilience weaknesses', 'design a safer recovery path'. Covers capacity under failure, partial completion, degraded behavior and recovery. Active incidents use incident-investigation; basic readiness checklists use service-lifecycle; change verdicts belong to reviewer.

root-causeskills/root-cause/SKILL.md

Use when diagnosing a bug, test failure, or unexpected behavior before permanent remediation, and especially after a fix attempt has already failed, or when guessing has started ("maybe it's X, let me try changing it"). Triggers: 'debug this failure', 'why did this test fail', 'the fix did not work'. Active response uses `incident-investigation`; a dispatched causal helper can use this method within its assignment. A bounded evidence lookup does not require a full diagnosis.

runbookskills/runbook/SKILL.md

Write or update an operational runbook: check/recover/verify/roll-back/escalate steps for one failure mode, the living-runbook accretion that grows it after every incident, and importing Confluence runbooks into the repo. Triggers: 'help me write a runbook for restarting pricing', 'update the runbook from this incident', 'import this Confluence runbook', 'apply the runbook structure'. Direct operational-document writing belongs to scribe, which selects runbook mode and applies this skill; retrospectives use postmortem.

service-lifecycleskills/service-lifecycle/SKILL.md

Take a service through its operational life on this stack: audit an existing service's readiness read-only, onboard a new or materially changed service, and retire one, each as a checklist a human executes under production-change-gate. Triggers: 'audit this service', 'onboard this service', 'retire this service', 'decommission this application'. Onboard and retire support draft planning; live execution requires an approved plan.

stack-profileskills/stack-profile/SKILL.md

The single stack-definition point — what this team runs today, which stacks it authors versus only supports, the stay-in-lane rule, and the platform boundary. Load before editing code in a support-only stack (Java/JVM), before recommending any runtime, tool, or infrastructure change, and when choosing between observability backends. Triggers: "what's our stack", "should we use X for this", "can we move this to Kubernetes / the cloud", "which backend do I query". This skill bundle changes when the ground shifts.

toil-reductionskills/toil-reduction/SKILL.md

Analyze recurring operational work and design measurable elimination, simplification, self-service or automation improvements. Triggers: 'reduce these repeated manual interventions', 'which toil should we automate', 'is this automation worth maintaining'. Uses frequency, effort, interruption, rework and maintenance evidence. Implementation belongs to software-engineer or the application owner; active incidents use incident-investigation.

Plugin manifests2

.claude-plugin/plugin.json
{
  "$schema": "https://json.schemastore.org/claude-code-plugin-manifest.json",
  "name": "save-toolkit",
  "displayName": "Save Toolkit",
  "description": "Application-engineering and site-reliability agents and reusable skills.",
  "version": "0.50.0",
  "author": {
    "name": "latent-sre",
    "url": "https://github.com/latent-sre"
  },
  "homepage": "https://github.com/latent-sre/save-toolkit",
  "repository": "https://github.com/latent-sre/save-toolkit",
  "license": "MIT",
  "keywords": [
    "agents",
    "skills",
    "sre",
    "observability",
    "pcf"
  ]
}
plugin.json
{
  "$schema": "https://agent-plugins.org/schemas/1.0.0/plugin.schema.json",
  "name": "save-toolkit",
  "description": "Application-engineering and site-reliability agents and reusable skills.",
  "version": "0.50.0",
  "author": {
    "name": "latent-sre",
    "url": "https://github.com/latent-sre"
  },
  "homepage": "https://github.com/latent-sre/save-toolkit",
  "repository": "https://github.com/latent-sre/save-toolkit",
  "license": "MIT",
  "keywords": [
    "agents",
    "skills",
    "sre",
    "observability",
    "pcf"
  ]
}

If you maintain this plugin, link to this source-backed listing from your README so users can review its manifest and indexed components.

[save-toolkit on Agent Plugins Marketplace](https://pluginsmp.com/plugins/save-toolkit)