Best for
- Replacing the nightly script that silently fails with something trustworthy
- Designing ingestion from sources you do not control
- Pipelines where a re-run today produces different numbers than yesterday
What you give it
- The sources, what the data must become, and who consumes it with what freshness
- Decisions on the trade-offs it surfaces (latency vs cost, completeness vs speed)
What you get back
- A pipeline design: stages, contracts between them, scheduling and dependencies
- Failure behaviour designed, not discovered: retries, idempotency, dead-letter paths, backfills
- Data-quality checks at the boundaries and freshness/completeness monitoring that alerts someone
How it works
- Maps sources honestly: formats, volumes, how each source misbehaves (late, duplicated, changed history).
- Designs stages with contracts: what each stage guarantees to the next, raw data kept so everything downstream is recomputable.
- Makes every step idempotent — re-running must never double-count or corrupt.
- Decides late and changed data explicitly: watermarks, upserts, reprocessing windows.
- Puts quality checks and freshness monitoring at the boundaries, wired to alert a human.
Example
You: We pull orders from three marketplaces into our warehouse nightly. It breaks weekly and nobody notices for days.
Result: A staged design: per-source ingestion with watermarks into raw tables (kept verbatim), idempotent transforms keyed on source ids, a reconciliation step that counts source vs warehouse per day, and alerts on freshness and completeness — plus a documented backfill procedure that re-runs any day safely.
Limits — please read
- It designs within your existing data stack; new platforms are a proposal with a case, not a default.
- Source systems you do not control will still surprise you; the design contains the blast radius, it cannot abolish it.
- Streaming vs batch is chosen from your freshness needs and budget — stated, with the cost of each.