ai-ml-testinglisted
Install: claude install-skill mejbaurbahar/fagun
# AI / ML Testing
## Scope note
This skill covers classical/predictive ML (classification, regression, ranking, recommendation models). For generative/LLM-specific concerns (prompt testing, hallucination, RAG), see [[llm-testing]]; for conversational agents built on top, see [[chatbot-voice-agent-testing]].
## Core evaluation metrics — use the right one for the problem
- **Accuracy** — fraction correct; misleading on imbalanced classes (99% accuracy predicting "not fraud" on a 1%-fraud dataset is a useless model).
- **Precision** — of predicted positives, how many were actually positive; matters when false positives are costly (flagging a legit transaction as fraud).
- **Recall** — of actual positives, how many were caught; matters when false negatives are costly (missing an actual fraud case).
- **F1** — harmonic mean of precision/recall, useful single number when both matter and classes are imbalanced.
- **AUC-ROC / AUC-PR** — threshold-independent view of separability; prefer PR curve over ROC when the positive class is rare.
Pick the metric based on the business cost of each error type — report multiple metrics, never just accuracy alone, and always ask what the actual cost asymmetry is if it's not stated.
## Model regression testing
- Maintain a fixed evaluation set (not the live/growing training data) and re-run it on every model version — a metric that improves in aggregate can still regress badly on a specific important slice.
- **Slice-based evaluation**: break do