Agents / Skills

Eval Dataset Builder

Skill

Builds the test set your AI feature is judged by: real cases collected and labelled, the hard and ugly ones deliberately included, stratified so the average cannot hide the failures — the ground truth that makes quality measurable.

Best for

  • AI features evaluated today on five cherry-picked examples
  • Building the regression set before the next prompt or model change
  • Turning 'users complained' into labelled cases that prevent recurrence

What you give it

  • The task, access to real inputs (logs, samples), and the judgement calls when labelling is genuinely ambiguous

What you get back

  • The dataset: real cases with expected outputs or rubric labels, each with provenance — where it came from, why it is here
  • Deliberate coverage: the strata that matter (input types, lengths, difficulties, failure-prone classes) each represented and tagged, so scores slice
  • The living-set apparatus: versioning, the intake path for new failures, and the contamination hygiene that keeps the test honest

How it works

  1. Collects from reality first: production inputs, support complaints, the logs — synthetic cases only to fill gaps reality has not yet supplied, and marked as synthetic.
  2. Stratifies deliberately: the axes along which failure differs (type, length, difficulty, sensitivity) each tagged, because the overall average is where failures hide.
  3. Oversamples the ugly: hard cases, past failures and edge classes beyond their natural rate — a set that mirrors traffic flatters the system on exactly the cases that burn trust.
  4. Labels with provenance and process: who judged, by what rubric, with ambiguous cases resolved explicitly — and keeps the set versioned, growing and uncontaminated.

Example

You: Build the eval set for our support-reply drafter; today we judge it by trying three tickets and squinting.

Result: The set: 240 cases from real (anonymised) tickets, stratified by the axes that matter — ticket type, length, tone (the angry ones oversampled on purpose), policy-sensitivity (every refund-adjacent case in), and the known failure classes from support's complaint log (each complaint became a labelled case). Labels: expected-content rubrics per case, with the 31 genuinely ambiguous ones resolved by the team's judgement calls, recorded. The apparatus: version 1 frozen, new production failures flow in through the intake path (labelled monthly), and the prompt that gets tuned never sees the set. The next prompt change: evaluated in minutes, two regressions caught in the angry-ticket stratum that the overall average had hidden.

Limits — please read

  • Labels inherit the labellers' judgement; genuinely contested cases get recorded as contested — a false gold standard measures compliance with a coin flip.
  • Personal data in real cases needs the anonymisation pass before the set exists anywhere permanent.
  • The set measures what it contains; it says which claims (novel inputs, adversarial users) remain outside its reach.