loom-model-evaluation

Solid

Evaluates ML models for performance, fairness, and reliability.

AI & Automation 56 stars 3 forks Updated today MIT

Install

View on GitHub

Quality Score: 87/100

Stars 20%
58
Recency 20%
100
Frontmatter 20%
70
Documentation 15%
100
Issue Health 10%
80
License 10%
100
Description 5%
100

Skill Content

# Model Evaluation ## Overview Choose metrics that match the business cost and data distribution, prevent leakage, validate with the right CV scheme, calibrate and threshold deliberately, and monitor for drift in production. This skill is the decision layer above sklearn/eval tooling. ## Metric selection (the highest-leverage decision) Wrong metric = confidently shipping a bad model. Pick from the cost structure and class balance, not habit. | Situation | Use | Avoid / why | | --- | --- | --- | | Rare positives (fraud, disease, churn) | **PR-AUC**, F-beta, MCC, recall@fixed-precision | **Accuracy** (a 99%-negative dataset scores 99% by predicting all-negative). **ROC-AUC** looks great even when precision is unusable — it ignores the huge TN base | | FN much costlier than FP | F-beta with β>1 (recall-weighted), recall@precision floor | plain F1 (β=1 assumes equal cost) | | FP costlier than FN | precision, F-beta β<1 | recall-optimized metrics | | Need a probability, not a label | log loss, **Brier score**, calibration curve | thresholded accuracy/F1 | | Multi-class imbalance | **macro** F1 (equal class weight), MCC | **micro**/weighted (dominated by majority class) | | Regression with outliers | MAE, median AE, Huber | MSE/RMSE (squares dominated by outliers) | | Regression, relative error matters | MAPE / SMAPE | RMSE; ⚠ MAPE explodes near zero and is asymmetric (penalizes over-prediction less) | | Ranking / retrieval | NDCG, MAP, MRR | accuracy | - **ROC-AUC vs PR-AUC:...

Details

Author
cosmix
Repository
cosmix/loom
Created
9 months ago
Last Updated
today
Language
Rust
License
MIT

Similar Skills

Semantically similar based on skill content — not just same category