data-engineering-master
Featured数据工程 — 数据平台从业者的认知操作系统, 覆盖把数据从源系统搬运成可靠 / 可查询 / 可信赖形态供分析 / ML / 数据产品消费的全生命周期 (生成 → 摄取 → 存储 → 转换 → 服务 + 安全/数据管理/DataOps/数据架构/编排/软件工程 六条暗流, Reis & Housley 框架): 摄取与集成 (批 + CDC 变更数据捕获 Debezium + EL 工具 Fivetran/Airbyte/Meltano/dlt + Kafka Connect + schema drift) / 存储与文件表格式 (对象存储数据湖 + 列存 Parquet/ORC/Arrow/Avro + 开放表格式 Apache Iceberg/Delta Lake/Apache Hudi + lakehouse + 分区/compaction) / 转换与建模 (ELT dbt/SQLMesh + Spark + 维度建模 Kimball + Inmon + Data Vault + 大宽表 OBT + 渐变维 SCD + 增量模型 + 语义/指标层) / 编排与工作流 (Apache Airflow/Dagster/Prefect/Mage/Kestra/Apache DolphinScheduler + DAG + 幂等 + 回填 backfill + 数据资产调度) / 批流与实时 (Apache Kafka/Apache Flink/Spark Structured Streaming/Kinesis/Pulsar/Redpanda + Lambda vs Kappa + watermark/窗口/exactly-once + 流式 SQL Materialize/RisingWave + 实时 OLAP ClickHouse/Apache Druid/Apache Pinot/StarRocks/Apache Doris) / 数仓与查询引擎 (Snowflake/BigQuery/Redshift/Databricks SQL/Trino/Presto/DuckDB/Polars + 存算分离 + MPP) / 数据质量测试与可观测性 (dbt tests/Great Expectations/Soda + 数据契约 + Monte Carlo data downtime + 新鲜度/量/schem
Install
Quality Score: 90/100
Skill Content
Details
- Author
- swaylq
- Repository
- swaylq/master-skill
- Created
- 4 months ago
- Last Updated
- 2 weeks ago
- Language
- Shell
- License
- MIT
Integrates with
Similar Skills
Semantically similar based on skill content — not just same category
data-engineer
Builds data pipelines, lakehouse architectures and data infrastructure. Use for implementation and operation. For pipeline design patterns and schema evolution, use data-pipeline-architect.
data-engineer
Data Engineer for Solaris - ELT/ETL pipeline design, orchestration (Airflow/Dagster/Prefect), dbt modeling standards (staging/intermediate/marts, tests, docs, semantic layer, Mesh), warehouse architecture and cost (Snowflake, BigQuery, Databricks), data quality gates and freshness SLAs, ingestion (CDC via Debezium/Fivetran/Airbyte, batch vs streaming with Kafka), schema evolution and data contracts, backfills. Use whenever the owner says "data pipeline", "ETL", "ELT", "Airflow", "Dagster", "dbt", "Snowflake", "BigQuery", "Databricks", "data warehouse", "data lake", "Kafka", "streaming", "CDC", "Fivetran", "Airbyte", "data quality", "freshness", "incremental model", "backfill", "warehouse cost", or asks why a pipeline is late, slow, expensive, or producing duplicates.
data-engineer
Handles data collection, ingestion, cleaning, and pipeline design