system-benchmark
FeaturedBenchmark a design system against named public systems (Material, Carbon, Polaris, GOV.UK) and maturity profiles. Triggers: benchmark our system, how do we compare, are we behind or ahead. Not for an internal health check (system-health).
Install
Quality Score: 90/100
Skill Content
Details
- Author
- murphytrueman
- Repository
- murphytrueman/design-system-ops
- Created
- 6 months ago
- Last Updated
- 2 days ago
- Language
- Python
- License
- MIT
Bundled in these plugins
Similar Skills
Semantically similar based on skill content — not just same category
system-health
Holistic health check across tokens, components, docs, adoption, governance, AI readiness, platform maturity, with status labels. Triggers: how healthy is my system, system health check, big picture. Not for one area (use its audit) or external comparison (system-benchmark).
benchmark
Local-only regression / benchmark skill for ui-clone-skills maintainers. Drives the standard ui-reverse-engineering pipeline against the canonical reference site (https://realfood.gov) and records AE/SSIM, iteration count, gate fail counts, and outcome to benchmark/history.csv so prompt / sub-doc / model-version drift surfaces as a trend. Trigger phrases: "run benchmark" / "regression benchmark" / "benchmark clone". The Makefile no longer has a `benchmark` target — setup is inline bash in this skill (Step 1 below). Internal: NOT registered in `.claude-plugin/plugin.json` `skills`. Not part of the public 3-skill marketplace surface. Maintainer tooling only.
skill-benchmark
Use when the user runs /skill-benchmark to score agent skills via LLM judges with baseline comparison, regression detection, and trend analysis, or to compare candidate models on a shared task set in a ranked table with per-model spend tracking. Not for release gating — use skill-benchmark-gate.