Best for
- Datasets guarded today by hope and a not-null constraint
- Pipelines that either fail on trivia or pass catastrophes
- Making 'is the data okay?' a dashboard instead of a debate
What you give it
- The dataset, what it means, and the past incidents you remember (profiling fills the gaps)
What you get back
- The rule set: per field and per dataset, each rule with its reason and its severity
- Placement and routing: which checkpoint runs which rule, and what each severity does — block, quarantine the row, or log the trend
- The aggregate sentinels: volume, freshness and distribution checks that catch the disasters row-level rules cannot see
How it works
- Derives rules from meaning and from history: what must be true for this data to be fit, and what has actually gone wrong before.
- Profiles the data to calibrate: real null rates, real value domains, real volumes — rules against fantasy baselines cry wolf by Thursday.
- Routes by severity because one red light teaches teams to ignore red: blockers stop loads, quarantines isolate rows with owners, warnings trend on dashboards.
- Adds the aggregate layer — volume, freshness, null-rate and distribution drift — where the silent disasters (the upstream change, the dropped partition) actually show.
Example
You: Write validation for the customer events feed; last month a schema change upstream silently nulled half the revenue column.
Result: The rule set: 23 row-level rules (keys unique and never null, amounts non-negative with the refund exception, enums closed against the documented legend, timestamps within sane bounds) each with severity routed — the 4 blockers stop the load, 11 quarantine the row with an owner, 8 log trends; plus the aggregate sentinels that would have caught last month in minutes: null-rate drift per critical column (revenue null-rate above 2% pages), volume bounds per batch, and value-distribution shift on amount. The feed's quality is now a dashboard, and the next upstream surprise was caught at 06:12, before a single report read it.
Limits — please read
- Rules encode known badness; the novel disaster needs the aggregate sentinels and a human — it sets both up.
- Over-validation is real: every rule costs runtime and attention; each must earn its place with a reason.
- Fixing violations upstream is the real win; the rules make the case with counts (a quality-steward agent pairs well).