Best for
- Grading a thousand outputs where humans can grade fifty
- Turning 'rate this 1-10' (which measures nothing) into criteria that measure something
- Knowing whether your judge agrees with humans before trusting its verdicts
What you give it
- What good output looks like (examples and judgement calls), and a sample of outputs for calibration
What you get back
- The rubric: quality decomposed into separately-judged criteria, each with anchored levels and worked examples — not one mushy score
- The judge prompt: engineered against the known biases (position, length, self-preference, sycophancy) with the mitigations that measurably matter
- The validation: judge-versus-human agreement on a labelled sample, reported per criterion — with the criteria the judge cannot be trusted on, named
How it works
- Decomposes quality into criteria a judge can actually hold: one dimension per judgement, anchored levels with worked examples — composite scores blur exactly the failures you need visible.
- Engineers the judge prompt against the documented biases: criterion isolation, evidence-before-verdict ordering, length and position controls, the source fenced from the output being judged.
- Validates before trusting: the judge grades a human-labelled sample; agreement is reported per criterion, and criteria below the bar get redefined or returned to humans.
- Monitors in production: sampled human audit and periodic re-validation, because judges drift with model versions and rubric edits.
Example
You: We need to grade 2,000 generated product descriptions weekly; human review covers 50.
Result: The rubric: five criteria judged separately (factual consistency with the source data, completeness of required elements, tone fit, length discipline, banned-claims compliance) — each with three anchored levels and a worked example per level, because 'rate quality 1-10' had produced judge scores that correlated with length and little else. The judge: one criterion per call (the all-at-once version measurably blurred), sources fenced, the bias mitigations applied. The validation: against 150 human-labelled cases — agreement strong on four criteria, weak on tone fit (humans disagreed with each other too; the criterion was redefined to concrete markers, re-validated). Production: the judge grades all 2,000 weekly with sampled human audit; the drift alarm (judge-human agreement re-checked monthly) has fired once, caught a rubric-version mismatch, and was worth the whole setup.
Limits — please read
- A judge inherits the rubric's clarity; criteria humans cannot agree on will not be rescued by a model — they get redefined to concrete markers or dropped.
- Judge agreement is measured against your labellers; if their standard shifts, re-validate.
- Some criteria (deep factual verification beyond the provided source) exceed judge reach; those stay human or get tooling.