← ClaudeAtlas

wargamelisted

Battle-test the toolkit — run a wargame scenario end-to-end against the skills, score the run with the rubric, and log gaps to fix. Use when the user says "wargame", "battle test", "run scenario N", or wants to stress-test the toolkit.
bedardandy/AI-appraisal-assist · ★ 0 · AI & Automation · score 65
Install: claude install-skill bedardandy/AI-appraisal-assist
# Wargame Run the toolkit against a fictional-but-realistic assignment that hides traps, then grade how well the skills caught them. ## Steps 1. **Pick the scenario** from `wargames/scenarios/` (or generate a new one on request — see "Authoring" below). Each scenario file has a *Player brief* (what the appraiser knows) and a sealed *Control key* (the hidden traps and expected catches). **Read only the Player brief during the run.** 2. **Run the pipeline blind, in a fresh context.** Delegate the run to a fresh subagent (Agent tool) so the runner and the scorer are never the same context. Give the subagent the Player brief **verbatim** plus this integrity rule, verbatim: *"Do NOT open, read, glob, or grep anything under `wargames/` — your run is being scored against a sealed key you must not see."* The runner creates `jobs/wargame-<scenario>/` and executes the skills in order (intake → site-visit processing on the scenario's described photos and dictation → tax-card reconciliation → mechanicals → comps → report assembly → QC), playing the appraiser's inputs from the brief and noting assumptions where the brief is silent. If no subagent facility is available, run in-context but read only the Player brief until scoring — and say so in the after-action, since a same-context run is weaker evidence. 3. **Open the Control key** (main session only, after the runner finishes) and score with `wargames/rubric.md`: - Traps caught / partially ca