← ClaudeAtlas

ai-ml-testinglisted

Use when testing a classical ML model or ML-powered feature — accuracy/precision/recall/F1 evaluation, model regression testing, bias/fairness testing, and ML pipeline robustness. For LLM/generative-AI-specific testing (prompts, hallucination, RAG), use llm-testing instead.
mejbaurbahar/fagun · ★ 1 · AI & Automation · score 72
Install: claude install-skill mejbaurbahar/fagun
# AI / ML Testing ## Scope note This skill covers classical/predictive ML (classification, regression, ranking, recommendation models). For generative/LLM-specific concerns (prompt testing, hallucination, RAG), see [[llm-testing]]; for conversational agents built on top, see [[chatbot-voice-agent-testing]]. ## Core evaluation metrics — use the right one for the problem - **Accuracy** — fraction correct; misleading on imbalanced classes (99% accuracy predicting "not fraud" on a 1%-fraud dataset is a useless model). - **Precision** — of predicted positives, how many were actually positive; matters when false positives are costly (flagging a legit transaction as fraud). - **Recall** — of actual positives, how many were caught; matters when false negatives are costly (missing an actual fraud case). - **F1** — harmonic mean of precision/recall, useful single number when both matter and classes are imbalanced. - **AUC-ROC / AUC-PR** — threshold-independent view of separability; prefer PR curve over ROC when the positive class is rare. Pick the metric based on the business cost of each error type — report multiple metrics, never just accuracy alone, and always ask what the actual cost asymmetry is if it's not stated. ## Model regression testing - Maintain a fixed evaluation set (not the live/growing training data) and re-run it on every model version — a metric that improves in aggregate can still regress badly on a specific important slice. - **Slice-based evaluation**: break do