Best for
- AI features judged today by 'it seems fine'
- Comparing models or prompts with numbers instead of anecdotes
- Catching quality regressions before customers do
What you give it
- The task your AI feature performs and examples of success and failure
- Judgement calls on the rubric: what counts as good, what is disqualifying
What you get back
- A test set built from real cases, labelled and version-controlled
- Metrics that fit the task — not one generic score hiding three failure modes
- A repeatable harness with a regression gate: change the prompt or model, see the delta
How it works
- Builds the test set from real traffic and real failures, stratified so the hard cases cannot hide in the average.
- Designs metrics per failure mode: correctness, faithfulness, format, tone — scored separately.
- Uses the right judge per metric: exact checks where possible, rubric-guided model grading where not, with the judge itself spot-checked against human labels.
- Runs repeatedly (variance is real) and reports distributions, not single lucky runs.
- Installs the harness as a gate: every prompt/model change runs it, deltas reported per metric and per case slice.
Example
You: We want to switch to a cheaper model for our document summariser but we are scared.
Result: A 150-case set from real documents (the long ones, the tables, the scanned mess included), a rubric scoring coverage, faithfulness and length separately, both models run three times each: the cheaper model matches on coverage, wins on length discipline, loses 4 points on faithfulness with tables — decision made with eyes open, and the harness stays as the regression gate.
Limits — please read
- An evaluation is only as good as its cases; it will keep pushing real, ugly data into the set.
- Model-graded metrics carry judge bias — it measures and reports that gap rather than pretending it away.
- Offline scores do not capture everything; it will tell you which claims need an online check with real users.