prompt-evaluatorlisted
Install: claude install-skill amitadar83/prompt-evaluator
# Prompt Evaluator
Evaluate whether a prompt helps with **its intended job**. A long, confident, or well-formatted prompt is not necessarily an effective prompt.
Review works in an Agent Skills-compatible assistant. Evidence tooling requires Python 3.10+. Live comparisons require an independently configured model runner or isolated agent sessions; none is bundled.
## Choose the requested deliverable
- **Review:** Inspect the prompt and concrete use case. Give evidence-backed weaknesses and targeted improvements. No model experiment is required. Label the result `Editorial review — not a benchmark`.
- **Design a test:** Create relevant cases, criteria, a baseline, and an execution plan. Label it `Test plan — not executed`.
- **Compare / run:** Run only when independent candidate and judge execution is available and authorized. Follow [references/benchmark.md](references/benchmark.md). If unavailable, deliver a runnable test plan or analyze supplied evidence; never substitute imagined answers for executions.
When “check this prompt” is ambiguous, start with a review, not a large experiment. Do not automatically rewrite or run the prompt when the user only asks for a diagnosis.
## Establish the evaluation contract
Identify the intended job, target model/environment, representative input, desired output, and costly failure modes from supplied context. Ask one focused question only if the missing information blocks a meaningful result. If proceeding with an assumption, stat