loom-model-evaluation
SolidEvaluates ML models for performance, fairness, and reliability.
Install
Quality Score: 87/100
Skill Content
Details
- Author
- cosmix
- Repository
- cosmix/loom
- Created
- 9 months ago
- Last Updated
- today
- Language
- Rust
- License
- MIT
Similar Skills
Semantically similar based on skill content — not just same category
ai-ml-testing
Use when testing a classical ML model or ML-powered feature — accuracy/precision/recall/F1 evaluation, model regression testing, bias/fairness testing, and ML pipeline robustness. For LLM/generative-AI-specific testing (prompts, hallucination, RAG), use llm-testing instead.
model-evaluation
Evaluation methodology — split strategy, metric selection, baseline comparison, failure mode analysis.
compare-llm-models
Use this to pick or switch the LLM behind a feature, based on evidence instead of hype or the newest release. Trigger on "which model should I use", "is GPT/Claude/Gemini/Llama better for this", "should I switch models", "can a cheaper model do this", "compare models for my use case". Evaluate on YOUR task, not on leaderboards alone.