ai-mlops

Featured

Operates production MLOps for ML, LLM, and agent systems. Use when designing deployment, monitoring, retraining, incident response, or GenAI security workflows.

AI & Automation 80 stars 17 forks Updated 1 weeks ago MIT

Install

View on GitHub

Quality Score: 89/100

Stars 20%
64
Recency 20%
90
Frontmatter 20%
70
Documentation 15%
100
Issue Health 10%
80
License 10%
100
Description 5%
100

Skill Content

# MLOps & LLMOps - Production Operations Hub **July 2026 posture:** version every changeable artifact, gate every release with a regression-eval suite in CI, instrument the whole path with OpenTelemetry (pin GenAI convention version — the spec now lives in its own repo and is still evolving), treat tool/RAG context as untrusted input, and ship rollback plus incident playbooks before launch. This skill is the execution hub for **operating AI systems in production**: - **Classical ML ops**: ingestion, registries, feature stores, drift, retraining, promotion - **LLMOps**: serving, prompt/config lifecycle, online evals, cost controls, safety gates - **Agent runtime ops**: tracing, tool governance, approval paths, MCP-aware telemetry, rollback - **Governance**: privacy, supply chain, auditability, AI Act readiness, safety incident handling Use this skill for **production architecture, release gates, monitoring, incidents, and governance**. Use adjacent skills for modelling, retrieval depth, agent design, or inference internals. ## When To Use This Skill Activate this skill when the user asks for: - Deploying an ML, LLM, RAG, or agent-backed system to production - Designing serving, batch, hybrid, or multi-region runtime architecture - Adding observability, drift detection, alerting, retraining, or release gates - Writing incident runbooks, rollback plans, or go/no-go checklists - Hardening an AI system against prompt injection, RAG poisoning, tool abuse, or data leakage - B...

Details

Author
vasilyu1983
Repository
vasilyu1983/AI-Agents-public
Created
9 months ago
Last Updated
1 weeks ago
Language
Python
License
MIT

Integrates with

Similar Skills

Semantically similar based on skill content — not just same category

AI & Automation Listed

mlops-engineer

Designs and hardens the infrastructure that carries models from training through production serving. Use when the user says "set up a model registry", "build the training pipeline", "deploy this model to production", or "/agent-collab:mlops-engineer." Also offer this proactively when a project trains or serves models but has no versioned artifacts, no promotion gate, or no monitoring for prediction quality.

0 Updated today
sumitake
AI & Automation Listed

senior-mlops-engineer

Use when operating the platform that trains, evaluates, deploys, serves, monitors, and retires ML models: building or reviewing training pipelines, model registries, feature stores, batch or online inference services, shadow and canary rollouts, drift detectors, model cards, retraining triggers, or model governance. Triggers: MLOps, model registry, feature store, training pipeline, model serving, batch inference, online inference, real time inference, model deployment, model monitoring, drift detector, shadow deployment, canary model, model card, governance, AI governance, lineage, model rollback, retraining, Tecton, Feast, MLflow, Kubeflow, Vertex AI, SageMaker, BentoML, KServe, Ray Serve, Triton, ONNX, model signing. Produces registry entries, feature contracts, rollout plans, drift configs, model cards, serving SLO sheets, retraining policies. Not for building the model itself, see senior-ml-engineer. Not for generic compute infra, see senior-devops-sre.

0 Updated 1 months ago
iamdemetris
AI & Automation Featured

ai-llm

Guides the LLM lifecycle from strategy to deployment. Use when planning, comparing, fine-tuning, migrating, or operating LLM systems.

80 Updated 1 weeks ago
vasilyu1983