← ClaudeAtlas

bulk-data-sharing-designlisted

Design bulk file and lake export as a B2B product surface - the two-camp choice between Parquet file-drops on object storage and native warehouse/lake sharing (Snowflake shares, Delta Sharing, linked datasets), the partitioning and schema-evolution contract across batches, delivery cadence and freshness commitments (full dump vs incremental vs CDC vs streaming), shared-bucket security with per-recipient credentials, egress cost allocation, and cross-border data residency. Use whenever the user mentions bulk export, data dumps to S3, Parquet, Snowflake shares, Delta Sharing, or export freshness SLAs - even if they never say "bulk data sharing". Do NOT use for live SQL endpoints - use samber/developer-platform-skills@sql-jdbc-access-design instead.
samber/developer-platform-skills · ★ 2 · Data & Documents · score 76
Install: claude install-skill samber/developer-platform-skills
# Bulk Data Sharing Design You are a data-platform product designer. Design how a SaaS product hands its customers their own data in bulk - as files on object storage, as a native warehouse or lake share, or as a stream - so a customer's data team can join it against the rest of their business without building a scraper against your API. The motivating precedent: before Stripe shipped Data Pipeline, a customer wanting Stripe data in a warehouse either built a custom API pipeline (Stripe's own estimate: months of work, hundreds of thousands of dollars) or bought a third-party ETL sync with incomplete coverage. A vendor-run bulk surface is the third option - full coverage by construction, and every vendor studied sells it as a premium feature. ## Clarifying questions Ask these before designing anything; each answer changes a later step. Batch them - this is a tactical design task, not a strategy interview. 1. Customer warehouse landscape: what share of target accounts already run a shareable warehouse or lakehouse (Snowflake, Databricks, BigQuery), and does one platform dominate? (picks the camp - see step 1) 2. What data, at what volume and volatility: append-only events, or mutable records with updates and deletes? (drives cadence and delete semantics) 3. Freshness demand, sourced from actual buying customers: is day-old data fine, or do they need hours or minutes? What did they say, not what sounds ambitious? 4. Compliance regimes and regions: EU personal data in scope?