evals
v0.3.1LLM evaluation methodology, eval-suite design, and guided practice around the Claude Code eval runner, distilled from Anthropic's official evaluation guidance: a knowledge router over success criteria, eval design, and grading methods (/evals:methodology); an action skill that interviews for measurable success criteria and scaffolds a graded eval suite for an LLM app or a Claude Code skill (/evals:design); a guided runner that preflights the CLI and the target, prices a suite before it spends, and reads the with-versus-without delta (/evals:plugin-eval); and a static case-file check that spends nothing and makes no model call (/evals:validate).
By Melodic SoftwareLicense: MIT20 GitHub starsUpdated 57 minutes ago
Directory evidence
- Runtimes
- Claude Code
- Parsed components
- 4 skill or MCP entries
- Source updated
- Sep 28, 2026
- Manifest status
- Canonical path parsed
The directory validates manifest shape and source location. It does not execute the plugin or provide a security endorsement. Review the indexing methodology →
Install evals for Claude Code
claude plugin marketplace add IchenDEV/agent-plugin-mkt
claude plugin marketplace update agent-plugin-marketplace
claude plugin install evals@agent-plugin-marketplacePaste and run these commands in a terminal with Claude Code. They add and refresh the PluginsMP catalog, then install this plugin.
The installer fetches third-party code from the source repository shown on this page. This directory validates manifest structure and source location, but does not perform a security audit; review the manifest, components, and source before installing.
Get the source manually
git clone https://github.com/melodic-software/claude-code-pluginsClone the source repository, then follow its setup instructions to add the plugin to a compatible client. The plugin root is plugins/evals/.
Plugin files
├── .claude-plugin/plugin.json├── skills/design/SKILL.md├── skills/methodology/SKILL.md├── skills/plugin-eval/SKILL.md└── skills/validate/SKILL.md
Included Skills4
Design an evaluation suite for an LLM-based application or a Claude Code skill: interview for measurable success criteria, pick a grading method per criterion, and scaffold a criteria doc plus eval cases into the consumer repo. Use when: 'design evals', 'create an eval suite', 'scaffold evals', 'write evals for my skill', 'define success criteria for this app', 'set up LLM testing', 'build a test set for my prompt'. Not for eval-design theory questions (use /evals:methodology), not for statically validating an existing evals.json (use /skill-quality:check validate-evals when installed), and not for running or scoring a suite, which is /evals:plugin-eval and the CLI it guides.
Answers LLM-evaluation design questions from Anthropic's official evaluation guidance: success criteria, eval-suite design, and grading methods for LLM-based applications and Claude Code skills. Use when: 'define success criteria', 'how do I eval this', 'LLM eval', 'measure prompt quality', 'LLM judge', 'model-graded eval', 'golden answer', 'grading rubric', 'eval grading method', 'exact match vs LLM-graded', 'how many eval cases', 'is my success criteria measurable'. Knowledge (WHY/WHAT of eval design), not a runner; for scaffolding a suite use /evals:design, and for running and scoring a plugin's suite against a no-plugin baseline use /evals:plugin-eval.
Guided practice around the `claude plugin eval` CLI, which runs and scores a plugin's eval suite. This skill does the rest: preflight (version floor, sandbox backend, target type), static validation with no model call, a printed cost estimate under the configured ceiling, the run itself, and the with-versus-without delta read correctly. Use when: 'run my plugin evals', 'plugin eval', 'evaluate this plugin', 'eval my skill', 'does my skill actually fire', 'what is the delta', 'read my eval results', 'aggregate-result.json', 'eval CI gate', 'can this machine run evals', 'how much will this eval cost'. Not for designing success criteria (use /evals:design), not for the skill-creator evals.json format (use /skill-quality:check validate-evals when the skill-quality plugin is installed), and not for CLAUDE.md or rules, which every run strips.
Statically validate a `claude plugin eval` suite (`prompt.md`, `case.yaml`, `graders/*.md`) before any run spends money. A standard-library Python script reports FAIL for what the binary rejects at load (unknown frontmatter key, unknown grader option, no grader, duplicate grader name, non-positive weight, out-of-range runs / max_turns / timeout_seconds, an env key outside EVAL_[A-Z0-9_]*, an unsupported schema_version major) and WARN for the documented authoring mistakes (target: files, inline (?i), judge-only graders, a gated tool in allowed_tools, file_exists in a read-only case). Use when: 'validate my eval cases', 'check my eval suite', 'will this suite load', 'lint case.yaml', 'check my graders', 'why did my case fail to load', or before paying for a run. Not for the skill-creator `evals/evals.json` format.
Plugin manifests1
{
"$schema": "https://json.schemastore.org/claude-code-plugin-manifest.json",
"name": "evals",
"version": "0.3.1",
"description": "LLM evaluation methodology, eval-suite design, and guided practice around the Claude Code eval runner, distilled from Anthropic's official evaluation guidance: a knowledge router over success criteria, eval design, and grading methods (/evals:methodology); an action skill that interviews for measurable success criteria and scaffolds a graded eval suite for an LLM app or a Claude Code skill (/evals:design); a guided runner that preflights the CLI and the target, prices a suite before it spends, and reads the with-versus-without delta (/evals:plugin-eval); and a static case-file check that spends nothing and makes no model call (/evals:validate).",
"author": {
"name": "Melodic Software",
"email": "[email protected]"
},
"license": "MIT",
"userConfig": {
"max_cost_usd": {
"type": "number",
"title": "Eval run cost ceiling (USD)",
"description": "Ceiling /evals:plugin-eval passes to the CLI as --max-cost-usd. It bounds one invocation, not a session's total, and a run that reaches it stops mid-suite with partial results. Raise it for a suite whose estimate exceeds it.",
"default": 5,
"min": 0
},
"unlimited_cost": {
"type": "boolean",
"title": "Run with no cost ceiling",
"description": "Drop --max-cost-usd from the invocation so a suite always runs to completion. The estimate is still printed before the run; only the ceiling goes away.",
"default": false
}
},
"keywords": [
"evals",
"evaluation",
"success-criteria",
"llm-judge",
"grading",
"rubric",
"testing",
"knowledge",
"skill"
]
}For maintainers
If you maintain this plugin, link to this source-backed listing from your README so users can review its manifest and indexed components.
[evals on Agent Plugins Marketplace](https://pluginsmp.com/plugins/evals)