data-lake-platform
FeaturedDesigns lakehouse platforms across Iceberg, Delta, Hudi, and Paimon. Use when choosing catalogs, CDC paths, query engines, governance, or cost controls.
Install
Quality Score: 89/100
Skill Content
Details
- Author
- vasilyu1983
- Repository
- vasilyu1983/AI-Agents-public
- Created
- 9 months ago
- Last Updated
- 1 weeks ago
- Language
- Python
- License
- MIT
Integrates with
Similar Skills
Semantically similar based on skill content — not just same category
data-engineering-guidelines
Use when designing, building, or reviewing data-engineering work — lakehouse table design (Delta/Iceberg, medallion, partitioning, compaction, schema evolution, MERGE/CDC), realtime and streaming pipelines (Kafka, Spark Structured Streaming, Flink, watermarks, exactly-once), data governance, AI/ML data support (feature stores, vector/RAG), and data quality, reliability, and cost.
data-streaming
Designs streaming platforms for Kafka, Flink, CDC, and lakehouse ingestion. Use when planning event backbones, CDC pipelines, schema governance, or real-time lakehouse delivery.
health-data-lake
When the user wants to design, build, or operate a clinical / healthcare data lake, lakehouse, or warehouse. Use when the user mentions "health data lake," "clinical data warehouse," "healthcare lakehouse," "OMOP," "PCORnet CDM," "Sentinel CDM," "i2b2," "CMS BCDA," "OHDSI," "ATLAS," "HADES," "Athena vocabulary," "FHIR Bulk Data," "$export," "flat FHIR," "SQL-on-FHIR," "Pathling," "Epic Clarity," "Caboodle," "Cerner Millennium ETL," "EMPI," "data quality dashboard," "Achilles," "tokenization vault," "Delta Lake," "Iceberg," "bronze/silver/gold," "Unity Catalog," "Lake Formation," "Snowflake healthcare," "Databricks Lakehouse for Healthcare," or "HIPAA-eligible warehouse." For analytics on top of curated data, see population-health-analytics or clinical-research. For raw FHIR API integration, see fhir-integration. For HL7 v2 ingestion specifics, see hl7-v2.