smithy
v0.14.1Full dev pipeline: research, planning, TDD implementation, persona review panels, root-cause debugging, and QA/stress/perf testing — never assumes, asks or recommends.
By Dyas Nuhakim SLicense: MIT0 GitHub starsUpdated 1 hour ago
Directory evidence
- Runtimes
- Codex and Claude Code
- Parsed components
- 20 skill or MCP entries
- Source updated
- Oct 2, 2026
- Manifest status
- Canonical path parsed
The directory validates manifest shape and source location. It does not execute the plugin or provide a security endorsement. Review the indexing methodology →
Install smithy for Codex and Claude Code
codex plugin marketplace add IchenDEV/agent-plugin-mkt
codex plugin marketplace upgrade agent-plugin-marketplace
codex plugin add smithy@agent-plugin-marketplacePaste and run these commands in a terminal with Codex. They add and refresh the PluginsMP catalog, then install this plugin.
The installer fetches third-party code from the source repository shown on this page. This directory validates manifest structure and source location, but does not perform a security audit; review the manifest, components, and source before installing.
Get the source manually
git clone https://github.com/dyasnurhakim/smithy-claudeClone the source repository, then follow its setup instructions to add the plugin to a compatible client. The repository root is the plugin root.
Plugin files
├── .codex-plugin/plugin.json├── .claude-plugin/plugin.json├── skills/anneal/SKILL.md├── skills/assay/SKILL.md├── skills/blueprint/SKILL.md├── skills/burnish/SKILL.md├── skills/calibrate/SKILL.md├── skills/commission/SKILL.md├── skills/forge/SKILL.md├── skills/guild/SKILL.md├── skills/handover/SKILL.md├── skills/hone/SKILL.md├── skills/inspect/SKILL.md├── skills/jig/SKILL.md├── skills/pattern/SKILL.md├── skills/proof/SKILL.md├── skills/ring-test/SKILL.md├── skills/smithy/SKILL.md├── skills/strike/SKILL.md├── skills/temper/SKILL.md├── skills/using-smithy/SKILL.md└── skills/wield/SKILL.md
Included Skills20
Debugging: reproduce → read-only root-cause analysis → the user approves the fix → regression test first (jigsmith) → ONE review. Triggers: 'anneal', 'debug', 'why is this broken', a failing test or run with an unknown cause.
Research → spec: explore the code, turn every assumption into a question or a recommendation, write spec.md. Building a feature starts here. Triggers: 'assay', 'research this', 'write a spec', 'new feature'.
Spec → plan: ≤8 tasks, each with a verify check, a persona pass, parallel batch markers with proof, and one self-contained brief per task. Works from a verbal request too (writes a mini spec). Triggers: 'blueprint', 'plan this', 'break into tasks'.
Design review and polish: screenshot the live local UI, judge it against DESIGN.md (or stated rules), score it, and send the fixes the user approves through strike — with before/after screenshots as proof. Triggers: 'burnish', 'polish the UI', 'design review'.
View or change smithy settings for ALL projects (global) or THIS project — model and effort per role, TDD settings, gates, review panel, where memory lives. Tests that a model works before saving it. Triggers: 'calibrate', 'smithy config', 'change the model for review'.
Write project test personas from the system's real user roles (code evidence + a short interview). They power per-persona QA in wield and the end-user view in guild. Triggers: 'commission', 'define the users', 'who uses this system'.
Build the plan task by task with forger or jigsmith agents (each self-checks and makes one commit), then ONE review of the whole job. Parallel batches in worktrees. Works without a plan for a single task. Triggers: 'forge', 'implement the plan', 'build this'.
Production-readiness panel: several persona reviewers run in parallel (masters judge the craft, patrons judge the experience) → one PRODUCTION_READY or NOT_READY verdict. Works on a job or on any diff. Triggers: 'guild', 'ready to ship?', 'review panel'.
Session handoff where every claim cites evidence, so the next session resumes with no re-discovery. Covers the active job, open lanes and a half-done TDD task. Triggers: 'handover', 'handoff', 'save session', 'wrap up'.
Performance: measure a baseline first, at least 3 runs and the median, find hot spots with profiler proof, give recommendations only (never edits code). Works inside a job or on its own. Triggers: 'hone', 'why is it slow', 'benchmark'.
Code review with two verdicts (does it meet the spec? is the code good?) — every finding has proof, a severity reason and a confidence. Reviews a branch, a commit range or uncommitted changes. Triggers: 'inspect', 'review this change', 'code review'.
TDD: write failing tests first, then the code, then ONE clean commit per task — the test-first order is proven by snapshots, not extra commits. Triggers: 'jig', 'TDD this', 'test-first', bug fixes with a repro.
Design creation: a direction grounded in the product's own world, shown as HTML previews, then tokens, states, motion and voice written to the project DESIGN.md. Triggers: 'pattern', 'design system', 'make it look good'.
Stress / load test a running service against limits the user sets (never invented), on a local or user-approved target only. Works inside a job or on its own. Triggers: 'proof', 'stress test', 'load test'.
Unit tests for the changed code, per the stack playbook: write the missing ones, run them, flag flaky ones. Works inside a job or on its own. Triggers: 'ring-test', 'unit test this', 'write unit tests'.
Full pipeline orchestrator (research → plan → build + review → panel → test) with approval gates and ledger resume. Triggers: 'run smithy', 'build end to end', 'full pipeline', resume smithy work.
Fix lane for small KNOWN changes and for review/QA findings: mini plan → one yes → bug items test-first (jigsmith), other items by the forger → targeted tests → ONE review → one report. No spec needed. Triggers: 'strike', 'quick fix', 'fix these items', 'fix these findings', 'small change'.
Full test pass: runs ring-test, wield, proof and hone (skipping, with the reason, any suite that cannot run) and gives one READY or NOT READY verdict. Works inside a job or on its own. Triggers: 'temper', 'test everything', 'full test pass'.
Skill router: which smithy skill to use when, what each one needs to run, the priority rules, and the thoughts that mean STOP. A short digest is injected each session; invoke this for the full router.
Functional QA done the way a user would: run real flows, screenshots required for UI, a 0-100 health score, per-persona flows, severity tiers. Works inside a job or on its own. Triggers: 'wield', 'QA this', 'does it work'.
Plugin manifests2
{
"name": "smithy",
"version": "0.14.1",
"description": "Full dev pipeline: research, planning, TDD implementation, persona review panels, root-cause debugging, and QA/stress/perf testing — never assumes, asks or recommends.",
"author": {
"name": "Dyas Nuhakim S",
"email": "[email protected]",
"url": "https://github.com/dyasnurhakim"
},
"homepage": "https://github.com/dyasnurhakim/smithy-claude",
"repository": "https://github.com/dyasnurhakim/smithy-claude",
"license": "MIT",
"keywords": [
"pipeline",
"planning",
"tdd",
"code-review",
"testing",
"debugging",
"personas",
"workflow"
],
"skills": "./skills/",
"hooks": {},
"interface": {
"displayName": "Smithy",
"shortDescription": "Blacksmith-themed dev pipeline: assay, blueprint, forge, inspect, anneal, temper",
"longDescription": "Use Smithy to take work from idea to tested code: research (assay), verify-annotated planning (blueprint), implementation with per-task review and optional TDD (forge/jig), multi-persona production-readiness panels (guild), root-cause debugging (anneal), and a full testing family (temper: unit/QA/stress/perf) — with per-project memory that survives session death and model routing per pipeline role (GPT-5.6 sol/terra/luna under Codex, Claude models under Claude Code).",
"developerName": "Dyas Nuhakim S",
"category": "Developer Tools",
"capabilities": [
"Interactive",
"Read",
"Write"
],
"defaultPrompt": [
"Take this feature from idea to tested code.",
"Fix these review findings.",
"Is this ready for production?"
],
"websiteURL": "https://github.com/dyasnurhakim/smithy-claude"
}
}{
"name": "smithy",
"description": "Full dev pipeline: research (assay), planning (blueprint), implementation (forge), review (inspect), debugging (anneal), testing (temper: ring-test/wield/proof/hone). Blacksmith-themed, memory-backed, tier-routed models. Never assumes — asks or recommends.",
"version": "0.14.1",
"author": {
"name": "Dyas Nuhakim S",
"email": "[email protected]"
},
"homepage": "https://github.com/dyasnurhakim/smithy-claude",
"repository": "https://github.com/dyasnurhakim/smithy-claude",
"license": "MIT",
"keywords": [
"pipeline",
"planning",
"code-review",
"testing",
"debugging",
"orchestration",
"model-routing"
],
"skills": [
"./skills/"
],
"commands": [
"./commands/"
]
}For maintainers
If you maintain this plugin, link to this source-backed listing from your README so users can review its manifest and indexed components.
[smithy on Agent Plugins Marketplace](https://pluginsmp.com/plugins/smithy)