Best for
- Reports nobody trusts because the numbers moved again
- Customer data full of duplicates, blanks and 'test test'
- Catching bad data at entry instead of at the board meeting
What you give it
- Access to the data and what it is supposed to mean
- Rulings on ambiguous cases: which duplicate wins, what counts as valid
What you get back
- A profile of what is actually in the tables — nulls, duplicates, impossible values, drift
- Written quality rules enforced where data enters, not discovered where it embarrasses
- Root-cause fixes: the forms, imports and processes that create the bad rows, named
How it works
- Profiles before promising: completeness, uniqueness, validity, consistency and freshness, measured per field that matters.
- Turns 'supposed to mean' into written, testable rules — including the definitional fights (what IS an active customer?) surfaced for settlement.
- Places checks where data enters and moves: form validation, import gates, pipeline boundary tests.
- Traces systematic badness upstream to the creating process, because cleaning without fixing the source is a subscription.
- Reports quality as trends per rule, so trust is measured rather than vibed.
Example
You: Sales says the customer count is 48,000. Finance says 51,500. Make the data trustworthy.
Result: Profile: 6,200 duplicate customers (three different spellings of the same company), 1,100 test records in production, two systems disagreeing on what 'active' means. The definitions written and agreed, duplicates merged by the agreed rule, a uniqueness check at entry, and the import that creates 80% of the duplicates fixed — counts now reconcile within 0.2%, difference explained.
Limits — please read
- Definitional conflicts (two departments, two meanings) are settled by humans; it frames the decision sharply.
- Historical cleanup at scale is staged and reversible, with the merge rules you approved — never bulk-guessed.
- Perfect data is not the goal; fit-for-purpose is, and the purpose defines the rules.