Agent Plugins Marketplace
← All plugins

evals

v0.3.1

LLM evaluation methodology, eval-suite design, and guided practice around the Claude Code eval runner, distilled from Anthropic's official evaluation guidance: a knowledge router over success criteria, eval design, and grading methods (/evals:methodology); an action skill that interviews for measurable success criteria and scaffolds a graded eval suite for an LLM app or a Claude Code skill (/evals:design); a guided runner that preflights the CLI and the target, prices a suite before it spends, and reads the with-versus-without delta (/evals:plugin-eval); and a static case-file check that spends nothing and makes no model call (/evals:validate).

Claude Code4 Skills

By Melodic SoftwareLicense: MIT20 GitHub starsUpdated 57 minutes ago

Directory evidence

Runtimes
Claude Code
Parsed components
4 skill or MCP entries
Source updated
Sep 28, 2026
Manifest status
Canonical path parsed

The directory validates manifest shape and source location. It does not execute the plugin or provide a security endorsement. Review the indexing methodology →

Install evals for Claude Code

Installs for the current user
claude plugin marketplace add IchenDEV/agent-plugin-mkt
claude plugin marketplace update agent-plugin-marketplace
claude plugin install evals@agent-plugin-marketplace

Paste and run these commands in a terminal with Claude Code. They add and refresh the PluginsMP catalog, then install this plugin.

The installer fetches third-party code from the source repository shown on this page. This directory validates manifest structure and source location, but does not perform a security audit; review the manifest, components, and source before installing.

Get the source manually
git clone https://github.com/melodic-software/claude-code-plugins

Clone the source repository, then follow its setup instructions to add the plugin to a compatible client. The plugin root is plugins/evals/.

Plugin files

plugins/evals/
├── .claude-plugin/plugin.json
├── skills/design/SKILL.md
├── skills/methodology/SKILL.md
├── skills/plugin-eval/SKILL.md
└── skills/validate/SKILL.md

Included Skills4

designskills/design/SKILL.md

Design an evaluation suite for an LLM-based application or a Claude Code skill: interview for measurable success criteria, pick a grading method per criterion, and scaffold a criteria doc plus eval cases into the consumer repo. Use when: 'design evals', 'create an eval suite', 'scaffold evals', 'write evals for my skill', 'define success criteria for this app', 'set up LLM testing', 'build a test set for my prompt'. Not for eval-design theory questions (use /evals:methodology), not for statically validating an existing evals.json (use /skill-quality:check validate-evals when installed), and not for running or scoring a suite, which is /evals:plugin-eval and the CLI it guides.

methodologyskills/methodology/SKILL.md

Answers LLM-evaluation design questions from Anthropic's official evaluation guidance: success criteria, eval-suite design, and grading methods for LLM-based applications and Claude Code skills. Use when: 'define success criteria', 'how do I eval this', 'LLM eval', 'measure prompt quality', 'LLM judge', 'model-graded eval', 'golden answer', 'grading rubric', 'eval grading method', 'exact match vs LLM-graded', 'how many eval cases', 'is my success criteria measurable'. Knowledge (WHY/WHAT of eval design), not a runner; for scaffolding a suite use /evals:design, and for running and scoring a plugin's suite against a no-plugin baseline use /evals:plugin-eval.

plugin-evalskills/plugin-eval/SKILL.md

Guided practice around the `claude plugin eval` CLI, which runs and scores a plugin's eval suite. This skill does the rest: preflight (version floor, sandbox backend, target type), static validation with no model call, a printed cost estimate under the configured ceiling, the run itself, and the with-versus-without delta read correctly. Use when: 'run my plugin evals', 'plugin eval', 'evaluate this plugin', 'eval my skill', 'does my skill actually fire', 'what is the delta', 'read my eval results', 'aggregate-result.json', 'eval CI gate', 'can this machine run evals', 'how much will this eval cost'. Not for designing success criteria (use /evals:design), not for the skill-creator evals.json format (use /skill-quality:check validate-evals when the skill-quality plugin is installed), and not for CLAUDE.md or rules, which every run strips.

validateskills/validate/SKILL.md

Statically validate a `claude plugin eval` suite (`prompt.md`, `case.yaml`, `graders/*.md`) before any run spends money. A standard-library Python script reports FAIL for what the binary rejects at load (unknown frontmatter key, unknown grader option, no grader, duplicate grader name, non-positive weight, out-of-range runs / max_turns / timeout_seconds, an env key outside EVAL_[A-Z0-9_]*, an unsupported schema_version major) and WARN for the documented authoring mistakes (target: files, inline (?i), judge-only graders, a gated tool in allowed_tools, file_exists in a read-only case). Use when: 'validate my eval cases', 'check my eval suite', 'will this suite load', 'lint case.yaml', 'check my graders', 'why did my case fail to load', or before paying for a run. Not for the skill-creator `evals/evals.json` format.

Plugin manifests1

plugins/evals/.claude-plugin/plugin.json
{
  "$schema": "https://json.schemastore.org/claude-code-plugin-manifest.json",
  "name": "evals",
  "version": "0.3.1",
  "description": "LLM evaluation methodology, eval-suite design, and guided practice around the Claude Code eval runner, distilled from Anthropic's official evaluation guidance: a knowledge router over success criteria, eval design, and grading methods (/evals:methodology); an action skill that interviews for measurable success criteria and scaffolds a graded eval suite for an LLM app or a Claude Code skill (/evals:design); a guided runner that preflights the CLI and the target, prices a suite before it spends, and reads the with-versus-without delta (/evals:plugin-eval); and a static case-file check that spends nothing and makes no model call (/evals:validate).",
  "author": {
    "name": "Melodic Software",
    "email": "[email protected]"
  },
  "license": "MIT",
  "userConfig": {
    "max_cost_usd": {
      "type": "number",
      "title": "Eval run cost ceiling (USD)",
      "description": "Ceiling /evals:plugin-eval passes to the CLI as --max-cost-usd. It bounds one invocation, not a session's total, and a run that reaches it stops mid-suite with partial results. Raise it for a suite whose estimate exceeds it.",
      "default": 5,
      "min": 0
    },
    "unlimited_cost": {
      "type": "boolean",
      "title": "Run with no cost ceiling",
      "description": "Drop --max-cost-usd from the invocation so a suite always runs to completion. The estimate is still printed before the run; only the ceiling goes away.",
      "default": false
    }
  },
  "keywords": [
    "evals",
    "evaluation",
    "success-criteria",
    "llm-judge",
    "grading",
    "rubric",
    "testing",
    "knowledge",
    "skill"
  ]
}

If you maintain this plugin, link to this source-backed listing from your README so users can review its manifest and indexed components.

[evals on Agent Plugins Marketplace](https://pluginsmp.com/plugins/evals)