trajectoryeval
v0.1.11Trajectory-match verifier sibling: checks the PATH a worker took (its ordered tool-call sequence) against an expected trajectory spec, in strict/unordered/subsequence modes. v0.1.7: adds a tolerance-based `fuzzy_hash` diff strategy — `TierConfig::threshold_permille` config plus a pure `fuzzy_diff()` computing a deterministic permille drift distance (differing leaves / total leaves) that tolerates drift at or under the threshold and escalates to `DiffOutcome::DriftedBeyondThreshold` above it; distance is stored as `u32` to keep the `Eq` derives on `DiffOutcome`/`TierVerdict` intact (no `f64`). v0.1.6: the seeded 1-in-N sampling hash now delegates to `harness_core::hash::fnv1a64` (the single canonical FNV-1a implementation) instead of a private reimplementation, closing a code-duplication finding; hash values are unchanged (same algorithm/constants), so sampling behavior is bit-for-bit identical. v0.1.5: `tier` subcommand is now e2e-tested end-to-end via a committed example core-flow allowlist (examples/tier-config.json); seeded 1-in-N non-core sampling has a rate-pinning test (catches subtler mutations than mere non-constancy); a core flow using the unimplemented `screenshot` diff strategy now yields an explicit tri-state `needs-human` verdict (distinct exit code 3) instead of silently masquerading as a hard diff failure. v0.1.4: adds risk-tiered e2e verification (`tier` subcommand) — a config-driven core allowlist where core flows get an every-run deterministic structured-data snapshot diff (perceptual-hash/screenshot behind a documented stub boundary) and non-core flows get an existence check or seeded low-frequency sampling. Subscription-native (one bundled Rust binary, no API key).
By yukineko0 GitHub starsUpdated yesterday
Directory evidence
- Runtimes
- Claude Code
- Parsed components
- 0 skill or MCP entries
- Source updated
- Sep 23, 2026
- Manifest status
- Canonical path parsed
The directory validates manifest shape and source location. It does not execute the plugin or provide a security endorsement. Review the indexing methodology →
Install trajectoryeval for Claude Code
claude plugin marketplace add IchenDEV/agent-plugin-mkt
claude plugin marketplace update agent-plugin-marketplace
claude plugin install trajectoryeval@agent-plugin-marketplacePaste and run these commands in a terminal with Claude Code. They add and refresh the PluginsMP catalog, then install this plugin.
The installer fetches third-party code from the source repository shown on this page. This directory validates manifest structure and source location, but does not perform a security audit; review the manifest, components, and source before installing.
Get the source manually
git clone https://github.com/yukineko/claude-harnessesClone the source repository, then follow its setup instructions to add the plugin to a compatible client. The plugin root is crates/trajectoryeval/.
Plugin files
└── .claude-plugin/plugin.json
Plugin manifests1
{
"name": "trajectoryeval",
"version": "0.1.11",
"description": "Trajectory-match verifier sibling: checks the PATH a worker took (its ordered tool-call sequence) against an expected trajectory spec, in strict/unordered/subsequence modes. v0.1.7: adds a tolerance-based `fuzzy_hash` diff strategy — `TierConfig::threshold_permille` config plus a pure `fuzzy_diff()` computing a deterministic permille drift distance (differing leaves / total leaves) that tolerates drift at or under the threshold and escalates to `DiffOutcome::DriftedBeyondThreshold` above it; distance is stored as `u32` to keep the `Eq` derives on `DiffOutcome`/`TierVerdict` intact (no `f64`). v0.1.6: the seeded 1-in-N sampling hash now delegates to `harness_core::hash::fnv1a64` (the single canonical FNV-1a implementation) instead of a private reimplementation, closing a code-duplication finding; hash values are unchanged (same algorithm/constants), so sampling behavior is bit-for-bit identical. v0.1.5: `tier` subcommand is now e2e-tested end-to-end via a committed example core-flow allowlist (examples/tier-config.json); seeded 1-in-N non-core sampling has a rate-pinning test (catches subtler mutations than mere non-constancy); a core flow using the unimplemented `screenshot` diff strategy now yields an explicit tri-state `needs-human` verdict (distinct exit code 3) instead of silently masquerading as a hard diff failure. v0.1.4: adds risk-tiered e2e verification (`tier` subcommand) — a config-driven core allowlist where core flows get an every-run deterministic structured-data snapshot diff (perceptual-hash/screenshot behind a documented stub boundary) and non-core flows get an existence check or seeded low-frequency sampling. Subscription-native (one bundled Rust binary, no API key).",
"author": {
"name": "yukineko"
},
"keywords": [
"trajectory",
"agentevals",
"verifier",
"llmops",
"rust"
]
}For maintainers
If you maintain this plugin, link to this source-backed listing from your README so users can review its manifest and indexed components.
[trajectoryeval on Agent Plugins Marketplace](https://pluginsmp.com/plugins/trajectoryeval)