Agent Plugins Marketplace

Recently added

Plugins grouped by the day they were added to the directory, newest first.

Sep 23, 2026

  1. md-eval0stars

    eval-docs skill: score Markdown docs for AI coding agent readiness and human readability with the md-eval CLI (TypeSafe Jev).

    Claude Code1 skill
  2. Scores each plan with md-eval before Claude leaves plan mode and sends weak plans back once with the gaps to fix.

    Claude Code
  3. narness0stars

    Harness engineering: constrain AI agents with code, hooks and scripts instead of prompts — skills, validation scripts, git hooks and per-tool configs

    Claude Code1 skill
  4. workflowy0stars

    Workflowy 노드 하나를 작업 흐름(workstream)의 기록으로 삼아, Claude Code 세션들이 그 아래에 작업 과정을 이어서 정리해 기록한다.

    Claude Code1 skill1 MCP server
  5. Two modes, one review page. Plan a change as pseudocode and get it approved before any code is written, or map an existing diff into the same diagram, with the real code attached, to review someone else's change. Feedback goes straight back to the agent.

    Claude Code2 skills
  6. Shape Up SDLC harness for Claude Code: shaping, intake, orient, scope-mapping, building (T0-verified, sandboxed, scope-contracted), evaluation and QA skills orchestrated by a tech-lead.

    Claude Code12 skills
  7. Google Search Console reports in Claude, Cursor and other MCP clients. Local and read-only.

    Claude Code1 MCP server
  8. Vaadin tools for AI agents: inspect and validate Vaadin projects with a self-contained native CLI.

    CodexClaude Code2 skills
  9. Personal agent skills for real engineering.

    Claude Code
  10. Remove multi-vendor AI provenance marks: invisible Unicode (Layer A), statistical text watermarks via rewrite (Layer B), and C2PA/EXIF/XMP/container metadata on images, documents, and media. Ships the remove-ai-marks skill (thin client over the repo's HTTP service) and the self-contained clean-user-facing-text skill.

    Claude Code
  11. bend-spec0stars

    Spec-first work for Bend projects traced by bolt (ez, bolt, eztoml, ezhttp, ezjson, snap, shake): bringing a project under spec, and adding features to one that is.

    Claude Code2 skills
  12. Agent skills for auditing and improving Agent Skills against the spec, Anthropic best practices, and ICM context-management criteria.

    CodexClaude Code2 skills
  13. Generate, match, and edit Neural DSP amp-sim presets from a song, reference recording, or plain-English change. Personal interoperability with your own licensed copy of the plugin.

    Claude Code3 skills
  14. プロダクトをエージェントに引き渡すときにだけ必要になる固有部品を持つプラグイン。法務ドラフト(利用規約・プライバシーポリシー・返金ポリシー)の雛形、サポート窓口メールの設定雛形、教訓ログの雛形の3つだけを配る。インフラ構築・開発ワークフロー導入・auto-merge 配線・SNS 運用はそれぞれ infra / dev-workflow / sns-autopilot が持つため本プラグインは呼び出さず、導入順だけを README に書く。住人側(workspace・cron・チャンネル)は flatmate の new-resident が担い、本プラグインはプロダクトのリポ側だけを扱う。

    Claude Code
  15. issue / PR 単位の API 換算コストを、Claude Code の会話ログ(`${CLAUDE_CONFIG_DIR:-$HOME/.claude}/projects/**/*.jsonl`)から集計する。帰属はブランチと(リポジトリ識別子, issue 番号)の 2 本立てで、サブエージェント(isSidechain)の消費も同じブランチに寄せる。`/cost <番号>` で PR ならヘッドブランチの総額、issue なら投稿を境界に切った区間の合計を USD と円で返す。pr-review-gate が `agent-review:passed` を付けた直後に、PostToolUse の hook がその PR へ `/cost` の 1 行目を 1 本のコメントで貼る(環境変数 `COST_LEDGER_GATE_REPORT=off` で止まる)。実行時に python3(3.8 以上)と、番号の判別に認証済みの gh を必要とする。git は 2.31 以上(rev-parse --path-format=absolute を使う)。

    Claude Code
  16. fusebase6stars

    Comprehensive FuseBase MCP server and agent platform tool suite for workspaces, Y.js live collaboration, dashboards, visual tables, CRM pipelines, client portals, automations, and Gate PostgreSQL stores.

    CodexClaude Code1 MCP server
  17. huitzo0stars

    Build Intelligence Packs and Dashboards on the Huitzo platform: SDK, manifest, CLI, platform API/MCP and dashboard SDK references, docs-first workflow skills, developer/reviewer agents, safety hooks and a project-docs MCP server.

    Claude Code
  18. Required workflow checks with Claude Code function hooks. Refuses a knowledge-file write until knowledge-save is open, refuses hand edits to generated indexes, refuses a pull request, close or merge until knowledge-save (and the work or merge-and-clean-up skill) is open, and holds a final reply once when working memory or the index rebuild and checker did not follow the change. Fact checks only: no model call and no reading of the agent's words. Runs only when CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1.

    Claude Code
  19. 萬達寵物通用 AI 工具包. Includes 7 skills.

    Claude Code7 skills
  20. Ragdoll 工作區插件. Includes 4 skills.

    Claude Code4 skills
  21. NorwegianForest 工作區插件. Includes 10 skills.

    Claude Code10 skills
  22. Maltese POS 工作區插件. Includes 6 skills.

    Claude Code6 skills
  23. Create and review survey drafts, publish requested studies, and analyze QuestionPunk results.

    Claude Code4 skills1 MCP server
  24. dev-lead2stars

    Cross-model delegation and adversarial review for CLI coding agents: one lead, six runtime adapters, and no change merges reviewed only by its own model family.

    Claude Code14 skills
  25. qa-kit0stars

    QA toolkit: clarify and solve problems (grill-me, brainstorm, problem-solving, predict), generate test cases and edge cases, write Playwright/Vitest/k6 automation, run and audit test suites, explore apps in a browser, and debug failures to root cause.

    CodexClaude Code9 skills
  26. Make any API agent-usable — even one with broken docs, no OpenAPI, a WAF, or an auth wall. Comprehend a spec (or JS-rendered docs) into first-call-correct MCP tools your agent calls right the first time, with auth injected and no integration code. On a painful real API, first-call-correctness goes from 10% to 65% (measured); wire x402 pay-per-call; defend against poisoned specs. Installs a live demo surface (TxODDS) over MCP — try it with zero setup.

    Claude Code1 MCP server
  27. Comprehend any API or Solana program into first-call-correct agent tools, and check the call before it counts: simulate the exact bytes, predict compute, refuse on stale evidence.

    Agent Plugins5 skills
  28. Buy from a Let Me Buy storefront on Solana and confirm delivery, with the traps that the program's IDL does not state.

    Agent Plugins2 skills
  29. Trajectory-match verifier sibling: checks the PATH a worker took (its ordered tool-call sequence) against an expected trajectory spec, in strict/unordered/subsequence modes. v0.1.7: adds a tolerance-based `fuzzy_hash` diff strategy — `TierConfig::threshold_permille` config plus a pure `fuzzy_diff()` computing a deterministic permille drift distance (differing leaves / total leaves) that tolerates drift at or under the threshold and escalates to `DiffOutcome::DriftedBeyondThreshold` above it; distance is stored as `u32` to keep the `Eq` derives on `DiffOutcome`/`TierVerdict` intact (no `f64`). v0.1.6: the seeded 1-in-N sampling hash now delegates to `harness_core::hash::fnv1a64` (the single canonical FNV-1a implementation) instead of a private reimplementation, closing a code-duplication finding; hash values are unchanged (same algorithm/constants), so sampling behavior is bit-for-bit identical. v0.1.5: `tier` subcommand is now e2e-tested end-to-end via a committed example core-flow allowlist (examples/tier-config.json); seeded 1-in-N non-core sampling has a rate-pinning test (catches subtler mutations than mere non-constancy); a core flow using the unimplemented `screenshot` diff strategy now yields an explicit tri-state `needs-human` verdict (distinct exit code 3) instead of silently masquerading as a hard diff failure. v0.1.4: adds risk-tiered e2e verification (`tier` subcommand) — a config-driven core allowlist where core flows get an every-run deterministic structured-data snapshot diff (perceptual-hash/screenshot behind a documented stub boundary) and non-core flows get an existence check or seeded low-frequency sampling. Subscription-native (one bundled Rust binary, no API key).

    Claude Code
  30. tracekit0stars

    Span-tree tracer for condukt runs: record one run's interpreter→worker→verifier phases as a parent-linked span tree (phase, model, ms, cost, status), render it with `tracekit trace <RID>`, and export OpenTelemetry GenAI-semconv JSON. Answers which phase of a failed run was slow/expensive/broke. Subscription-native (one bundled Rust binary, file-only, no network, no API key).

    Claude Code
  31. tdd0stars

    v0.1.27 integrates this line with the remote's: the session-scoped skip now carries a third state. Session-scoping fixed WHOSE token is consumed; it did not fix WHEN. A Stop is adjudicated by four independent gate processes, so a token honoured by one that allows was burned even when another blocked the same stop — the stop never happened and the one-shot escape had been spent on nothing, which is standing pressure toward a permanent bypass (CLAUDE.md 5). `consume_session_skip` now takes `stop_hook_active` and parks an honoured token as `<id>.skip.honoured`: it is re-honoured while a re-entry after a block is in flight, and deleted the first time a chain actually ends. A rename failure that can neither bound nor clear the token is resolved to NOT honouring it. Test-first gate for Claude Code: a Stop hook that blocks the turn when implementation lines land and no test is visible in the uncommitted changes (it inspects the working tree only, never already-committed tests), plus a /tdd skill and red/green/verify subcommands that make test-first (RED before GREEN) a verifiable artifact. Subscription-native (one hook + skill + bundled Rust binary, no API key). v0.1.25: the shared project-root `.tdd-skip` marker is gone. The operator escape is now `tdd skip --reason "..."`, which is scoped to the issuing session (CLAUDE.md 5 forbids a shared one-shot file under parallel sessions, since whichever session stops next consumes it), requires a non-empty reason, and appends both the issue and the consumption to the gate log so a bypass cannot happen unrecorded. v0.1.20: config.rs's local GATE_CRATES const is now `pub use harness_core::fleet::GATE_CRATES;`, removing a duplicate hand-written copy that had independently drifted (lost `overwatch`) in the past. v0.1.18: gate verdict now routes through harness_core::verdict (compass DoD1).

    Claude Code1 skill
  32. taskprog0stars

    Multi-session progress file for Claude Code. Injects .claude/progress.md at SessionStart so the agent knows what is pending; at Stop it prompts the agent to keep the file current. Enables seamless HOTL handoff across sessions.

    Claude Code1 skill
  33. Stuck-loop detector + escalation for Claude Code: a PostToolUse hook that spots repeated identical actions and edit thrash, then injects an escalating nudge to change approach or ask the user for help. Subscription-native (one hook + bundled Rust binary, no API key).

    Claude Code
  34. specguard0stars

    v0.2.52: shard 監査 subagent の「並列で同時に起動してよい」を最大 3 体ずつの波に変更 (run と spec-audit の両方)。shard は間引かない。 仕様↔実装 整合監査ハーネスを Claude Code から subscription-native に実行する。各 shard を read-only な in-session subagent で監査し (nested claude --print なし)、決定的ハーネス (scope/render/parse/report) は specguard バイナリに委譲する。

    Claude Code
  35. ship0stars

    v0.1.9: cargo fmt --all reformatting only (fixes the 'build & commit plugin binaries' smoke workflow's `cargo fmt --all --check` gate, chronically red on main since 2026-07-21 per scripts/check-ci-red.py); no behavior change. nudges the commit・merge・push・plugin-update shipping ritual; detects unshipped git/plugin-cache state; subscription-native

    Claude Code1 skill
  36. Per-session work metrics for Claude Code: PostToolUse + Stop hooks roll up tool calls, turns, files touched, a size class (XS–XL) and a work category per session, viewable with `session-insights report` and optionally logged as a dated note to an Obsidian vault. The `/record` command (`record-now`) regenerates the note's machine-owned 数値サマリ/コスト blocks and has the model author Japanese prose sections (完了サマリ/つまずき・学び/振り返り/注意点・落とし穴/残課題/要追跡・あとで確認/関連); cross-session backlog reconciliation now delegates to the standalone `backlog` crate (`~/.backlog/tasks.toml`) rather than a session-insights-owned vault backlog.md. Subscription-native (two hooks + bundled Rust binary, no API key).

    Claude Code
  37. scout0stars

    v0.1.5: 5 レンズを 1 メッセージで全部並列起動していたのを最大 3 体ずつの波に変更。レンズは間引かない (上限は波を分ける理由であって調査範囲を削る理由ではない)。 Multi-lens project audit that generates actionable tasks (施策) for Claude Code: a /scout skill that gathers deterministic project state, fans out read-only sub-agents across five lenses (current issues, security, industry/peer-project practices via web search, missing measures, safety), then dedupes/scores the findings into prioritized tasks, writes them to the backlog, and hands execution to /flow. Complements compass (single-goal gradient) as a broad-reconnaissance SOURCE. Subscription-native (skill only, no binary, no API key).

    Claude Code1 skill
  38. Schema-validation gate for LLM structured outputs at source→executor boundaries: validates named declared schemas, emits structured errors for one re-ask, and counts rejects to metrics so silent drops become observable. Subscription-native (one bundled Rust binary, no API key).

    Claude Code
  39. runbook0stars

    v0.1.8: cargo fmt --all reformatting only (fixes the 'build & commit plugin binaries' smoke workflow's `cargo fmt --all --check` gate, chronically red on main since 2026-07-21 per scripts/check-ci-red.py); no behavior change. Reusable procedure includes for Claude Code: a UserPromptSubmit hook that expands `!name` macros in your prompt into the matching repo-committed procedure (.runbook/<name>.md), so recurring workflows run the same way every time. Inspired by Devin Playbooks; distinct from knowledge injection. Subscription-native (one hook + bundled Rust binary, no API key).

    Claude Code
  40. v0.1.24 integrates this line with the remote's. The review is now scoped to what THIS session edited: `harness_core::transcript::files_edited_by_session` reads the session transcript and `review::attribute` narrows the changed-file list to it, so a peer session's uncommitted work in a shared tree is no longer presented as your diff. Attribution is three-valued — when the transcript cannot be read the list is NOT narrowed and the block says so, rather than silently reviewing nothing. Also: the session-scoped skip now carries a third state. Session-scoping fixed WHOSE token is consumed; it did not fix WHEN. A Stop is adjudicated by four independent gate processes, so a token honoured by one that allows was burned even when another blocked the same stop — the stop never happened and the one-shot escape had been spent on nothing, which is standing pressure toward a permanent bypass (CLAUDE.md 5). `consume_session_skip` now takes `stop_hook_active` and parks an honoured token as `<id>.skip.honoured`: it is re-honoured while a re-entry after a block is in flight, and deleted the first time a chain actually ends. A rename failure that can neither bound nor clear the token is resolved to NOT honouring it. Code-review gate for Claude Code: on Stop, review the diff before the agent can declare done. In inject mode it blocks the stop and injects a review rubric so the running agent self-reviews (no API key); in subprocess mode it runs an independent reviewer and surfaces only its findings. Subscription-native, bundled Rust binary. v0.1.21: the shared project-root `.reviewgate-skip` marker is gone. The operator escape is now `reviewgate skip --reason "..."`, which is scoped to the issuing session (CLAUDE.md 5 forbids a shared one-shot file under parallel sessions, since whichever session stops next consumes it), requires a non-empty reason, and appends both the issue and the consumption to the gate log so a bypass cannot happen unrecorded.

    Claude Code
  41. replaykit0stars

    Trace→golden replay regression harness: convert tracekit-recorded condukt run traces into evalkit golden replay cases. The sibling of curate (playbook→golden) — three subcommands (extract/verify/promote) pin a run's phase set, error count, and cost as a portable, self-verifying snapshot so regressions surface as a failing golden. Subscription-native (one bundled Rust binary, no API key).

    Claude Code
  42. propguard0stars

    v0.1.44 integrates this line with the remote's: the session-scoped skip now carries a third state. Session-scoping fixed WHOSE token is consumed; it did not fix WHEN. A Stop is adjudicated by four independent gate processes, so a token honoured by one that allows was burned even when another blocked the same stop — the stop never happened and the one-shot escape had been spent on nothing, which is standing pressure toward a permanent bypass (CLAUDE.md 5). `consume_session_skip` now takes `stop_hook_active` and parks an honoured token as `<id>.skip.honoured`: it is re-honoured while a re-entry after a block is in flight, and deleted the first time a chain actually ends. A rename failure that can neither bound nor clear the token is resolved to NOT honouring it. v0.1.41: the shared project-root `.propguard-skip` marker is gone. The operator escape is now `propguard skip --reason "..."`, which is scoped to the issuing session (CLAUDE.md 5 forbids a shared one-shot file under parallel sessions, since whichever session stops next consumes it), requires a non-empty reason, and appends both the issue and the consumption to the gate log so a bypass cannot happen unrecorded. v0.1.38: closes the fail-open v0.1.37 documented but deliberately left open (backlog 87dbfbb8 p0 + d8e22b26 p1). `run_git` now answers `Determination<String>` and `diff_text` answers `Determination<DiffText>`, forwarding (never re-minting) the boundary's `Undetermined` from any of the four reads that feed it: `git diff`, `git diff --cached`, the untracked `ls-files`, and an untracked file's body. `evaluate` maps that to a new `decide_diff_failed` — tag `diff-read-failed`, bounded by max_attempts, escapable, no hash recorded, no per-property violations attributed — placed BEFORE the `match cfg.mode` split so both inject and subprocess modes route through it. Measured before the fix with the real `git`: a tracked file whose working-tree content is not valid UTF-8 (no NUL, so git emits a textual diff) gave changed_files=Files(["bad.rs"]) with diff_text="" and truncated:false, and the gate returned ALLOW tag=empty-diff; the valid-UTF-8 control gave BLOCK tag=below-threshold. The partial case is closed the same way and deliberately returns no diff at all rather than the readable half: previously `git diff` succeeding while `git diff --cached` could not be decoded produced a diff mentioning a.rs, omitting b.rs, and carrying truncated:false — announcing itself complete over propguard's only incompleteness signal. `empty-diff` survives but now means only what it says: every read succeeded and the diff really was empty. New tests pin both faults end-to-end through `evaluate`, each with an anti-vacuity control that differs only in whether the bytes decode, plus a genuinely-clean-repo allow and a measured demonstration that Mode::Subprocess really can reach properties-satisfied (so the guard prevents a reachable allow, not a hypothetical one). v0.1.37: docs-only correction of a claim v0.1.36 itself introduced. `run_git`'s comment said dropping `None` matched a benign 'fail gracefully / treat as git-unavailable' convention because 'callers already tolerate None/empty output'. That is false on the decision path: `changed_files` recovers (None -> collect false -> ChangeScan::Failed -> block), but `run_diff` and the untracked `ls-files` read inside `diff_text` still drop `None` silently, and a tracked file whose working-tree content is not valid UTF-8 yields changed_files = Files(["bad.rs"]) with diff_text = "" and truncated:false, which gate.rs's `if diff.trim().is_empty() { allow("empty-diff") }` ALLOWS (the valid-UTF-8 control gives a 131-byte diff and a BLOCK). The comment now states which caller is hardened and which two are not, and names the open backlog ids 87dbfbb8 and d8e22b26. The behaviour is deliberately NOT fixed here. v0.1.36: an unreadable git stdout no longer reads as a clean repo — harness-core 0.2.3's bounded pipe read stopped discarding the read error and stopped folding an expired read budget into an empty string, so a `git` that exits 0 while its output never arrives now reaches ChangeScan::Failed (gate fails closed) instead of the Files(vec![]) a genuinely clean repo produces. v0.1.32: fixed a stale never-break-a-turn docstring/comment in main.rs and README.md — panics no longer swallow-to-exit-0; the Stop hook's panic path resolves via harness_core::gate::run::run_guarded's fail-closed policy (block on the first stop, bounded allow only on a second consecutive stop_hook_active panic). No behavior change. v0.1.16: merged two independently-developed Continuous-Audit fix lines that landed on the same v0.1.14 base — (a) stopped correlation-store pollution (CA-propguard-03/04): a checker-unavailable or diff-truncated Block evaluated no property, so it no longer stuffs the full derived prop_ids into the property_id-keyed fleet-correlation signal as if they were real per-property violations (it reports an empty set, mirroring the below-threshold narrowing); (b) documented the checker subprocess timeout (checker_timeout_secs, default 300s, process-tree kill on timeout) in README.md and added config-propagation tests covering the propguard.toml override path. Earlier v0.1.13 fixed 3 subprocess-hang gaps (CA-propguard-004/005/006) — checker timeout now kills the whole shell-spawned process tree via a Unix process group instead of just the direct shell child; stdout is read on a bounded thread so a lingering process holding the pipe open can't hang past the timeout; and all git subprocess calls now go through a single timeout-bounded choke point that fails gracefully instead of hanging the Stop hook indefinitely. Property gate for Claude Code: on Stop, derive 3–5 semantic properties (invariants) from a task's done_criteria and check the generated code against them before the agent can declare done. In inject mode it blocks and injects a property checklist for the running agent to self-verify (no API key); in subprocess mode it runs an independent checker and counts per-property PASS/FAIL, blocking when fewer than a threshold hold. Goes beyond 'concrete tests pass'. Subscription-native, bundled Rust binary.

    Claude Code
  43. Config-driven pre-commit static audit for Claude Code: a Stop hook that runs generic checks plus your project rules (from a TOML file) over the pending diff — shelling out to git and optional linters — and blocks the stop, feeding findings back to the agent until they're clean. Subscription-native (one hook + bundled Rust binary, no API key). v0.1.16: the shared in-tree `<audit_dir>/.audit-skip` marker is gone — it was the fifth mirror of the one-shot bypass the four Stop gates dropped in the same change. It carried no attribution and lived in the working tree, so whichever invocation ran next spent it: a bypass one session armed was routinely consumed by a different session's commit, or by a human's terminal `git commit`. The replacement is `precommit-audit skip --reason "..."`, scoped to the issuing session via CLAUDE_CODE_SESSION_ID, reason-required, and recorded at both issue and consumption. An invocation with no session id consumes nothing at all.

    Claude Code
  44. playbook0stars

    v0.1.7: cargo fmt --all reformatting only (fixes the 'build & commit plugin binaries' smoke workflow's `cargo fmt --all --check` gate, chronically red on main since 2026-07-21 per scripts/check-ci-red.py); no behavior change. Project knowledge retrieval + injection for Claude Code: a UserPromptSubmit hook that pulls the curated atomic notes relevant to your prompt and injects them under a strict char budget, so conventions and gotchas resurface without re-typing. Subscription-native (one hook + bundled Rust binary, no API key).

    Claude Code
  45. Per-session concurrency cap for Claude Code. A PreToolUse hook counts the Bash calls and subagents actually in flight for this session and denies the call that would exceed the cap (3 shells and 3 subagents by default); PostToolUse gives the slot back and the ledger is cleared at every turn boundary. Enforcement lives in a bundled binary, not in prose a model may ignore. Subscription-native (no API key).

    Claude Code
  46. overwatch0stars

    v0.2.31: v0.2.29 の finder 同時起動制限の散文を revert した。同時実行数の上限は parallelguard が PreToolUse で deny して強制する。あわせて origin/main を統合し、`skills/continuous-audit/SKILL.md` の既定 target list から taintguard (2026-08-24 にユーザー裁定で repo から撤去) を外した版を取り込んだうえで parallelguard を含める。したがって正典 GATE set は 7 (blastguard propguard specguard stuckguard mutategate overwatch parallelguard) で、`scripts/continuous-audit.sh` の `DEFAULT_TARGETS` および `scripts/rollout-plugins.sh` の `GATE_CRATES` と同期する (`check-gate-crates-sync.py` が機械照合)。skill prose のみ、コード変更なし。 v0.2.29: continuous-audit の finder 同時起動を最大 3 体に制限し、Step 2 の verifier と合わせて 3 を超えないことを明記。対象 crate は間引かない。 v0.2.28: SessionStart/Stop の `status` が 4 source (backlog / hypothesis / condukt / compass) すべてを bare 名で spawn していたため、hook プロセスに plugin の bin dir が PATH に無い環境では 4 本同時に `(unknown: No such file or directory (os error 2))` へ落ちていた。実測 2026-08-21、測定点 cd2576bd: claude プロセスの /proc/<pid>/environ に plugins/cache/yukineko は 0 件 — plugin bin dir の PATH 追加は Bash tool の shell 内だけで、hook はそれを継承しない。bare 名の spawn は ~/.cargo/bin に残っていた 2026-07-23 版の stale コピーが login PATH 上にあったために偶然動いていただけで、それを (正しく) 削除した 2026-08-20 の bb046648 以降 banner は全滅していた。condukt と compass は ~/.cargo/bin に一度も存在しなかったので、0.2.26 の三値化以前は `(none)` として無言で fail-open していた (hook 経路から一度も観測できていなかった)。修正は新設の `harness_core::plugin_bin::resolve` で plugin cache を第一候補・PATH を fallback として解決する — 順序は意図的に autoflow の既存 resolver の逆で、rollout が版を保証する唯一の配布経路であり、PATH 先行こそ 91fa24df の stale shadow を勝たせた原因だったため。F→P オラクル: `env PATH=/usr/local/bin:/usr/bin:/bin overwatch status` は修正前が 4 unknown (出荷済み SessionStart banner と逐語一致)、修正後は backlog pending 355 ほか全 source を報告する。v0.2.27 (merge reconciliation, no new code): two branches independently shipped DIFFERENT content as 0.2.26 -- the launcher exit-0 fail-open fix and the `overwatch status` tri-state fix (aggregate.rs/render.rs) -- so the label 0.2.26 ambiguously named two trees. This release is the union of both, renumbered so the version identifies one tree again. No behaviour beyond the two merged changes. v0.2.26: `overwatch status` (the SessionStart+Stop hook) rendered `(none)` for a source it could not read, identical to a source it read and found empty. Measured against the shipped binary: a truncated `leases.json` holding one LIVE lease from another session produced byte-identical output, exit 0, to a store that had never been written — `store::load_leases` already separated absent (`Ok(empty)`) from corrupt/unreadable (`Err`) via `boundary::read_to_string`'s `Determination`, but `aggregate::build` bound it with `if let Ok(..)` and threw the distinction away; the same collapse applied to the four subprocess sources (`shell_soft`'s `Option<String>` folded not-installed / non-zero-exit / non-UTF-8 into one `None`) and to the JSON/TSV parsers (`Err(_) => Default::default()`, `unwrap_or(0)`). `(none)` in the Sessions pane is the claim "no other session is live" — the fact CLAUDE.md §8 says never to assume, and the liveness input condukt's main-tree guard reads before permitting a commit in main's shared working tree (`condukt/src/maintree.rs` documents this exact flattening as "a real residual hole, not a safe degradation"). `ProgressView` now carries `undetermined: Vec<UndeterminedSource>`; unreadable sources render `(unknown: <reason>)` plus a loud stderr WARNING, and `status --json` gains a machine-readable `undetermined` key (omitted on the clean path, so the existing contract is unchanged). Parsers return `Result`; three tests that asserted garbage input yields `pending: 0` / all-zero buckets — writing the fail-open down as the contract — now assert `Err`. v0.2.0 (major bump, not content): 0.1.52 added pub fields to externally-constructible structs (`AuditRound.unverified`, `RoundMetric.unverified`, `AuditMetrics.cumulative_unverified`, `ReviewFinding.verdict`) without a breaking version bump, so `cargo semver-checks` correctly failed CI on main (17 consecutive runs) with `constructible_struct_adds_field`; this release only bumps 0.1→0.2 per the crate's own pre-1.0 convention (breaking change bumps the minor field) to match the version to the API shape already shipped — no code change. v0.1.48: `store::mark_branch_merged` (branch-keyed sibling of `mark_changeset_merged` — the merge path knows only the branch, not the `task_key`) plus `store::clear_runtime_overlap_holds` give condukt the on-land cleanup that takes a merged branch's `ActualChangeset` out of the mid-flight overlap-detection set (fixing spurious ~30-min false-positive merge HOLDs where a cleanly-landed peer stayed `merged=false` within the lease TTL) and clears any stale `RuntimeOverlap` hold recorded against the reused branch name; `prune_stale_changesets` is now wired opportunistically on land to keep `active_changesets.json` bounded. All fail-soft under `LeaseLock`, with a regression test proving a landed branch is excluded from detection and its entry is pruned. v0.1.44: `lease::begin()`'s load->is_held_by_other check->save read-modify-write was unprotected against a TOCTOU race (two sessions racing begin() for the same key could both pass the check before either saved, both believing they'd claimed the lease) -- condukt::lock/backlog::lock had already fixed this exact pattern via hardlink+create_new(O_EXCL) exclusive locking, but overwatch::lease hadn't (hypothesis 9c733d74). New `crates/overwatch/src/lock.rs` (`LeaseLock`) ports that design -- hard-link atomic publish, a TMP_SEQ intra-process collision guard, stale-lock reap via pid liveness, bounded wait, fail-soft degrade-to-unlocked on timeout -- and `begin()` now holds it across the whole load->check->save cycle. New `crates/overwatch/tests/lease_concurrency.rs` spawns two real processes racing `begin()` for the same key across 8 trials and asserts exactly one wins; an env-gated `OVERWATCH_TEST_BEGIN_DELAY_MS` widens the race window past normal process-spawn overhead so the test deterministically forces the interleave (verified by temporarily reverting the lock: the test then reliably catches the double-claim). Also fixed a related pre-existing test-suite flake: `aggregate.rs` and `store.rs` each sandboxed the process-global `$HOME` env var behind their OWN separate `Mutex`, so tests in the two modules could still race each other's `HOME` mutation under parallel `cargo test`; unified to one crate-wide `store::HOME_ENV_LOCK`. v0.1.43: `review-metrics` now reports `stale_undisposed_with_fix_commit` -- a read-only recount (same commit-range/store logic as `reconcile-fixed`, but never writes) of findings whose fix commit has already landed but are still undisposed, printed as a WARNING line in the human-readable report too. Closes the last piece of the 2026-07-17 stale-review-queue gap: `reconcile-fixed` only clears the backlog when someone remembers to run it, so this makes the gap visible on every `review-metrics` call even if reconcile-fixed hasn't run that round. v0.1.42: new `reconcile-fixed` command closes the "fix commit landed, nobody ran record-disposition" gap that let review-queue go stale (2026-07-17 incident: 18 already-fixed findings sat "open" for weeks because nobody remembered to dispose them). Scans a range of git commit messages (`--since-ref`/`--range`/`--last-n`, default last 50) for `CA-<crate>-<NNN>` finding-id references and auto-records a CONFIRMED disposition (`auto-reconcile(commit <hash>)`) for any referenced finding present on the review-findings store but not yet dispositioned; idempotent (already-disposed ids are skipped, a finding referenced by multiple commits is disposed once). Fail-soft end-to-end: a missing git binary, a non-repo cwd, or a non-zero `git log` all degrade to "0 processed" rather than erroring, so it is always safe to wire into automation. Wired into `scripts/continuous-audit.sh` (via the existing fail-soft `run_ow` wrapper) so every Continuous-Audit round auto-reconciles already-fixed findings before bridging the remainder to the backlog. v0.1.39: `test_freshness::run_ignored_test`'s `cargo test` subprocess call used `Command::output()`, which waits unbounded — a hung/deadlocking `#[ignore]`d regression test could wedge `overwatch review-queue --to-backlog`'s per-finding loop forever. Switched to `spawn()` + `wait_timeout()` (60s, since the call includes a build step) mirroring `ctxrot::hooks::guard::run_with_timeout`; timeout kills and reaps the child, folding into the existing `ExecutionError` fail-soft path (`bridge.rs` unchanged). v0.1.35: fixed CA-overwatch-004 — the RECORD path (`overwatch audit-round record --confirmed N --regression-tests-added M`) constructed an `AuditRound` via `AuditRound::new` with NO clamp on `regression_tests_added`, unlike the separate CLOSE path (`set_round_tests`, which already clamps to `confirmed`); `record --confirmed 1 --regression-tests-added 999` previously persisted an unclamped 999 and propagated a closure_rate far above the documented [0,1] range. `AuditRound::new` now clamps `regression_tests_added` to `confirmed` for every construction path, and the `record` CLI's printed JSON reports the stored (clamped) value rather than the raw arg. v0.1.34: continuous-audit SKILL.md's "対象 crate (既定)" section understated the target list as 5 crates (blastguard/propguard/specguard/stuckguard/mutategate), omitting overwatch, even though scripts/continuous-audit.sh's actual DEFAULT_TARGETS has included overwatch since it joined GATE_CRATES (558f864) -- fixed the doc to match the real 6-crate default (docs/fix-gate-crates-drift.md). v0.1.32: fixed set_round_tests() so an over-large `--tests` count passed to `audit-round close` clamps to the round's own `confirmed` count before being stored, preventing `closure_rate` from exceeding the documented [0,1] range. v0.1.31: Continuous-Audit finding triage (docs/DESIGN-continuous-audit-triage.md) — `ReviewFinding` gains an optional `rationale` field (verifier's CONFIRMED-判定根拠, `#[serde(default)]` so pre-existing `review_findings.jsonl` rows keep reading) and `record-finding --rationale` wires it through; new `test_freshness` module reverse-looks-up a `#[ignore = "<finding-id>: ..."]` regression test by finding-id and re-runs it (`cargo test -p <crate> -- --ignored <fn>`), fail-soft (`ExecutionError`/`NotFound` on any cargo/crate trouble); `review-queue --to-backlog` now thickens each backlog task's notes with elapsed days since confirmation, the rationale (if any), and the regression-test freshness verdict (FAIL/PASS/no test), all advisory-only (core dedup/idempotency unchanged). v0.1.27: SessionStart+Stop `status` hook now shares a short-lived (10s TTL) on-disk cache (`aggregate::build_cached`) so the two hooks firing close together (end of one turn, start of the next) collapse into one full aggregate scan (~5 subprocess spawns + lease-store scan) instead of two; a cold/expired/corrupt cache always falls through to a fresh build (fail-soft, no observability lost — just bounded staleness). Also fixes a latent JSON round-trip bug: `ProgressView`'s `sessions`/`runs` (`Vec`) and `BacklogSummary`'s `pending_by_priority` (`BTreeMap`) were missing `#[serde(default)]` alongside their `skip_serializing_if`, so deserializing an omitted-when-empty field previously errored instead of reconstructing the empty collection. v0.1.25: strengthened `disposition_metrics` integration tests to patch ONLY the timestamp fields (`ts`/`resolved_ts`) of the CLI-written JSONL ledgers, keeping the CLI-serialized `finding_id`/`verdict`/`reviewer` intact, so `false_positive_rate`/`agreement_rate`/`by_verdict` assertions now exercise `record-disposition`'s real field serialization end-to-end (test-only; no behavior change). v0.1.24: new `compact-findings` command performs non-lossy compaction/rotation of the append-only `review_findings.jsonl` hot store: finding records whose finding_id has been resolved (bridged to the backlog, or dispositioned by a human) are MOVED (never deleted) into a cold `review_findings_archive.jsonl`, so `review-queue`'s human-surface read stays bounded to OPEN items rather than lifetime append volume; `review-metrics`'s median-latency join now reads hot plus archive so the metric never regresses after compaction. Atomic temp+rename rewrites (archive written before hot for crash-safety); idempotent (a run with no newly-resolved findings is a byte-identical no-op). v0.1.23: new `auto-approved` companion command (read-only, foreign-file bridge to condukt's `gate-decisions.jsonl`, fail-soft by path) surfaces the DENOMINATOR — the population of decisions condukt self-answered without a human — as a count plus a deterministic seeded sample (`--since`/`--sample`/`--seed`/`--json`), so a human can judge whether spot-check sampling coverage of mass auto-approval is adequate. v0.1.22: review-queue gains a fourth source — condukt's durable escalation queue (`escalate.rs`) is bridged in as `EntryKind::Escalation` (High severity), read fail-soft by path (no condukt crate dependency, so the harness-core<-overwatch<-blastguard<-condukt direction stays acyclic) so every human-awaiting item (blocked/GATED tasks, not just gate-check findings) shows in the one unified pane. v0.1.21: review-effectiveness measurement — new `record-disposition` (confirmed|dismissed|false-positive per finding_id) + `review-metrics` commands close the loop on the human review queue: a fail-soft `dispositions.jsonl` ledger joined against `review_findings.jsonl` yields false-positive rate, human-agreement rate, and median resolution latency (JSON or human-readable). v0.1.20: review-queue noise collapse — AI findings dedup by a content fingerprint (source+file+summary), not just exact finding_id, so independent reports of the same issue under different ids collapse to one row; repeated same-plugin rollbacks likewise collapse to one row. Both carry an additive `occurrences` count (and a `(Nx)` summary marker) so recurring noise doesn't flood the human review surface. v0.1.19: new `store::record_finding` library entry lets condukt's gate-check Escalate branch auto-populate the review-queue ai-finding stream (previously producer-less), so needs-human/gated verdicts surface on the risk-ranked human queue automatically. v0.1.18: review-queue now risk-ranks by normalized severity (High-first) then recency, so a stale high-severity item is never buried/evicted below fresh low-severity noise; --limit yields the top-K riskiest and reports the deferred remainder. Project-global cross-session execution ledger + dedup guard + PDO progress view + fleet-level correlated gate-violation detection. Manages claim registry with heartbeat-based liveness, skip-on-duplicate contract for distributed session coordination, and normalized violation-signature recurrence/escalation across tasks and sessions.

    Claude Code1 skill
  47. v0.1.13: `skills/hypothesis/SKILL.md` に YAML frontmatter を追加した。Claude Code は frontmatter の `name`/`description` から skill を登録するので、散文で始まるファイルは黙って読み飛ばされる — この skill は「壊れていた」のではなく最初から **存在しなかった** (`/hypothesis:hypothesis` が not found)。plugin 側からは不可視な失敗で、ファイルは在り、`git status` は clean、rollout は成功を報告する (CLAUDE.md 第3節: 「登録されなかった」と「登録するものが無い」が同じ出力に潰れる)。実測 2026-08-26 (測定点 58c779af) では全 39 crate の SKILL.md のうち 1 行目が `---` でないのはこの 1 件だけだった。新規テスト `tests/skills_frontmatter.rs` が frontmatter の存在・`name`/`description` の有無・`name` とディレクトリ名の一致を pin する (追加前に RED を観測済み。skills が 0 件のときに空虚に pass しないよう anti-vacuity assert 付き)。挙動の変更は無い。 v0.1.8: the store's load->mutate->save cycle is now protected by an exclusive file lock (`StoreLock`, ported from backlog::lock/condukt::lock::RunLock/overwatch::lock::LeaseLock) so two concurrent processes mutating the store (e.g. `validate`/`reject`/`confidence`) can no longer lose one update to a last-writer-wins race; a multi-process regression test proves the fix (and that it fails without the lock). PDO hypothesis lifecycle management — create, validate, and track discovery hypotheses aligned with compass goals.

    Claude Code2 skills
  48. v0.1.11: cargo fmt --all reformatting only (fixes the 'build & commit plugin binaries' smoke workflow's `cargo fmt --all --check` gate, chronically red on main since 2026-07-21 per scripts/check-ci-red.py); no behavior change. Unified HOTL status dashboard for Claude Code. A /status command runs a bundled Rust binary that aggregates budgetguard's spend ledger, gauge's session records (recent sessions + cost), and taskprog's progress file into one human-on-the-loop view. Subscription-native (binary + command, no API key). A lightweight SessionStart hook additionally warns (fail-soft, silent when healthy) if a registered hook's binary is missing from disk, or if a stray PATH binary (e.g. a stale ~/.cargo/bin copy) shadows a plugin-cache copy of the same name.

    Claude Code
  49. gauge0stars

    Local LLMOps telemetry for Claude Code: a Stop hook that reads each session's transcript and records token usage, cache hits, tool calls, latency, and estimated cost to a local store — then `gauge report` rolls it up by project, model, and day. `gauge subagents` attributes cost to each individual sub-agent (Task) live from its transcript, so callers like condukt can record true per-task cost instead of a lumped session total. v0.3.9: `gauge subagents --json` now also emits `tokens_input`/`tokens_output` per sub-agent, so callers can record real token usage alongside cost. Observability for your own agent runs, subscription-native (one hook + bundled Rust binary, no API key, nothing leaves the machine).

    Claude Code
  50. v0.1.28: fixes the defect v0.1.27's wiring fix exposed — the hook was reachable again, but `sync` could not finish from the store's own steady state. `cmd_sync` pulled BEFORE committing local appends, and pulled with `--ff-only`. Both abort in normal operation: the store files live inside the sync dir, so every `record` leaves them as uncommitted working-tree modifications (`git pull` then refuses with "Your local changes to the following files would be overwritten by merge", exit 1), and any second machine pushing makes the histories diverge (`--ff-only` then exits 128, "Not possible to fast-forward"). Both were reproduced against a copy of the real store. The order is now commit → pull `--no-rebase --no-edit` → push, with the push skipped only when the branch can be *established* to be level with its upstream (an undeterminable position pushes anyway). Reordering alone was still not enough: two machines appending to the same JSONL land their additions adjacent at end-of-file, which the default merge driver reports as a content conflict, so `sync` now ensures the sync dir's `.gitattributes` declares `episodes.jsonl`/`playbooks.jsonl` as `merge=union` — it keeps BOTH sides' lines, and the store is already deduplicated by content hash, so a duplicate line is recoverable where a dropped episode is not. Finally, a failed sync no longer fails into silence: its only caller is a `SessionEnd` hook whose exit code and stderr reach nobody, so the failure is written to `~/.fugu-router/sync-error.json` and surfaced by this plugin's own `UserPromptSubmit` hook (as `additionalContext` for the model AND `systemMessage` for the user) until a later sync succeeds and clears it. An unreadable or unparseable marker still produces a notice — its presence already says the last sync failed. New test `sync_ordering.rs` (4 tests, RED observed before GREEN) covers the steady state, idempotency on repeat, surfacing, and clearing. v0.1.27: `hooks/hooks.json` now declares the record-store sync as a `SessionEnd` hook (`${CLAUDE_PLUGIN_ROOT}/bin/fugu-router sync`, timeout 30). The wiring previously existed only as an orphaned absolute path in the user's `~/.claude/settings.json` (`~/.cargo/bin/fugu-router sync`), whose build had been renamed away — so the hook exited 127 every session with neither exit code nor stderr reaching the agent or the user, making a sync that had nothing to do indistinguishable from a sync that never ran (CLAUDE.md §1/§3: it went dark, not red). The damage was silent and cumulative — measured 2026-08-26: the last `fugu-router sync` commit in `~/.fugu-router/record-repo` is 2026-07-23 09:57:38 +0900, while `episodes.jsonl`/`playbooks.jsonl` in that repo kept growing as uncommitted, unpushed working-tree modifications for ~34 days, so every episode recorded in that window existed only on one machine. New test `session_end_sync_hook_wired.rs` pins the structural fix: the declaration must live in-plugin, every declared command must resolve through `${CLAUDE_PLUGIN_ROOT}` and never through `~/.cargo/bin`, and every declared subcommand must actually exist in the binary (RED observed before GREEN). fugu-style per-model routing for Claude Code orchestration. Learns from past task outcomes (which model passed verification, at what cost) and picks the cheapest Claude tier that historically clears similar work. Feeds condukt's suggested_model deterministically; records outcomes back to a local episode store. No API key, no embedding service — lexical k-NN over a JSONL store. v0.1.24: new mode axis `--mode fast|normal|high` for `route`/`suggest` — a deterministic clamp applied AFTER the policy has already picked a worker/verifier pair (never touches `decide`/`decide_bandit`'s learning logic). `fast` shifts the worker one tier down (capped at sonnet — opus never selected as worker or verifier); `high` shifts one tier up (capped at opus); `normal` is the backward-compatible identity default. The verifier is always recomputed from the clamped worker via `policy::verifier_model`, additionally capped at sonnet under `fast`. Precedence: `--mode` flag > env `FUGU_ROUTER_MODE` > config.toml `mode` > `normal`; an invalid env/config value is an explicit error + non-zero exit, never silently coerced to `normal`. A `gated` decision is a no-op under every mode (mirrors `downgrade_for_budget`'s own gated no-op). Ordering: the mode clamp runs BEFORE `downgrade_for_budget` — budget is a hard resource limit and wins over a mode preference, and the rationale records both the mode's shift and any budget negation so a `high` pick never silently disappears. `Episode` gains a measurement-only `mode: Option<String>` field (`record --mode`); an absent value means "not recorded", never conflated with `Some("normal")`. v0.1.20: new `duration-outliers` command flags models whose avg duration within a task class is a relative outlier vs other models in the same class (>1.5x the cross-model mean by default), and cross-references outlier-vs-normal effective pass rate — the measurement tool for PDO hypothesis ae64db03. v0.1.19: `Episode` gains measurement-only `route_basis`/`route_confidence`/`route_rationale` (the routing `Decision`'s provenance), `lines_added`/`lines_removed`, and `tokens_input`/`tokens_output`, all `Option<_>` and backward-compatible (older JSONL lines parse with these as `None`); `record` gains matching `--route-basis`/`--route-confidence`/`--route-rationale`/`--lines-added`/`--lines-removed`/`--tokens-input`/`--tokens-output` flags. None of this is consulted by `policy::route`/`decide_bandit` — it exists so routing decisions and task cost can be retrospectively correlated against actual pass/fail outcomes. v0.1.12: `store::append_playbook` now writes body+newline in one `write_all` call (mirroring `append`'s existing single-syscall pattern) instead of `writeln!`'s two syscalls, closing the same O_APPEND interleaving hazard for playbook records; a JSON-serialization failure now propagates as an `io::Result` error instead of silently writing an empty line. New regression test `concurrent_append_playbook_never_interleaves_records`. v0.1.13: every git subprocess in `cmd_sync` (pull/clone/status/add/diff/commit/push) is now bounded by a timeout (30s network ops, 10s local ops) via `wait-timeout`, killing a hung/stalled git process instead of wedging the sync command indefinitely. v0.1.14: `Episode` gains a measurement-only `duration_secs: f64` field (`#[serde(default)]`, backward-compatible with older JSONL lines) and `record` gains a `--duration` flag threading it through; not consulted by `policy::route`/`decide_bandit` — routing/scoring behavior is unchanged.

    Claude Code1 skill