Best for
- AI features evaluated today on five cherry-picked examples
- Building the regression set before the next prompt or model change
- Turning 'users complained' into labelled cases that prevent recurrence
What you give it
- The task, access to real inputs (logs, samples), and the judgement calls when labelling is genuinely ambiguous
What you get back
- The dataset: real cases with expected outputs or rubric labels, each with provenance — where it came from, why it is here
- Deliberate coverage: the strata that matter (input types, lengths, difficulties, failure-prone classes) each represented and tagged, so scores slice
- The living-set apparatus: versioning, the intake path for new failures, and the contamination hygiene that keeps the test honest
How it works
- Collects from reality first: production inputs, support complaints, the logs — synthetic cases only to fill gaps reality has not yet supplied, and marked as synthetic.
- Stratifies deliberately: the axes along which failure differs (type, length, difficulty, sensitivity) each tagged, because the overall average is where failures hide.
- Oversamples the ugly: hard cases, past failures and edge classes beyond their natural rate — a set that mirrors traffic flatters the system on exactly the cases that burn trust.
- Labels with provenance and process: who judged, by what rubric, with ambiguous cases resolved explicitly — and keeps the set versioned, growing and uncontaminated.
Example
You: Build the eval set for our support-reply drafter; today we judge it by trying three tickets and squinting.
Result: The set: 240 cases from real (anonymised) tickets, stratified by the axes that matter — ticket type, length, tone (the angry ones oversampled on purpose), policy-sensitivity (every refund-adjacent case in), and the known failure classes from support's complaint log (each complaint became a labelled case). Labels: expected-content rubrics per case, with the 31 genuinely ambiguous ones resolved by the team's judgement calls, recorded. The apparatus: version 1 frozen, new production failures flow in through the intake path (labelled monthly), and the prompt that gets tuned never sees the set. The next prompt change: evaluated in minutes, two regressions caught in the angry-ticket stratum that the overall average had hidden.
Limits — please read
- Labels inherit the labellers' judgement; genuinely contested cases get recorded as contested — a false gold standard measures compliance with a coin flip.
- Personal data in real cases needs the anonymisation pass before the set exists anywhere permanent.
- The set measures what it contains; it says which claims (novel inputs, adversarial users) remain outside its reach.