review-training-data-qualitylisted
Install: claude install-skill bastos/skills
# Review Training Data Quality
Test whether the dataset teaches the intended capability rather than merely producing an easy loss curve.
## Inspect deterministic distributions
Summarize every product-relevant axis by split and overall: source group, task lane, goal, format, label, abstention, difficulty, candidate count, sequence length, terminal state, role, and preferred-answer position.
Use the generic JSONL summary helper:
```sh
python scripts/summarize_jsonl_fields.py corpus.jsonl \
--field split --field lane --field goal --field label \
--output quality-distributions.json
```
Look beyond equal row counts. Verify group-safe splits, distinct source groups, reasonable joint distributions, and enough examples at safety boundaries. Flag any category whose dominance would let the model ignore important context.
## Review candidate and label quality
Check that:
- positives are legal, plausible, and supported by supplied context;
- hard negatives are tempting but wrong for an explainable reason;
- multiple acceptable answers are preserved when evidence supports them;
- abstention is available and labeled only when warranted;
- teacher outputs use supplied identifiers and validate without silent repair;
- prompts exclude reference answers and teacher-only metadata;
- every lane has sufficient facts to make a defensible choice.
Run a small balanced teacher preflight before labeling the full corpus. Stop on identifier, replay, legality, terminal-boundary, context, or