Agents / Skills

Data Validation Rules

Skill

Writes the validation rules a dataset actually needs: derived from what the data means and how it breaks, placed at the right checkpoints, with severities that route — block, quarantine, or warn — instead of one giant red light.

Best for

  • Datasets guarded today by hope and a not-null constraint
  • Pipelines that either fail on trivia or pass catastrophes
  • Making 'is the data okay?' a dashboard instead of a debate

What you give it

  • The dataset, what it means, and the past incidents you remember (profiling fills the gaps)

What you get back

  • The rule set: per field and per dataset, each rule with its reason and its severity
  • Placement and routing: which checkpoint runs which rule, and what each severity does — block, quarantine the row, or log the trend
  • The aggregate sentinels: volume, freshness and distribution checks that catch the disasters row-level rules cannot see

How it works

  1. Derives rules from meaning and from history: what must be true for this data to be fit, and what has actually gone wrong before.
  2. Profiles the data to calibrate: real null rates, real value domains, real volumes — rules against fantasy baselines cry wolf by Thursday.
  3. Routes by severity because one red light teaches teams to ignore red: blockers stop loads, quarantines isolate rows with owners, warnings trend on dashboards.
  4. Adds the aggregate layer — volume, freshness, null-rate and distribution drift — where the silent disasters (the upstream change, the dropped partition) actually show.

Example

You: Write validation for the customer events feed; last month a schema change upstream silently nulled half the revenue column.

Result: The rule set: 23 row-level rules (keys unique and never null, amounts non-negative with the refund exception, enums closed against the documented legend, timestamps within sane bounds) each with severity routed — the 4 blockers stop the load, 11 quarantine the row with an owner, 8 log trends; plus the aggregate sentinels that would have caught last month in minutes: null-rate drift per critical column (revenue null-rate above 2% pages), volume bounds per batch, and value-distribution shift on amount. The feed's quality is now a dashboard, and the next upstream surprise was caught at 06:12, before a single report read it.

Limits — please read

  • Rules encode known badness; the novel disaster needs the aggregate sentinels and a human — it sets both up.
  • Over-validation is real: every rule costs runtime and attention; each must earn its place with a reason.
  • Fixing violations upstream is the real win; the rules make the case with counts (a quality-steward agent pairs well).