Agents / Agents

Model Evaluator

Agent

Builds the evaluation harness that tells you whether your AI feature is actually good: real test sets, metrics that match what users care about, and regression gates so quality cannot silently slide.

Best for

  • AI features judged today by 'it seems fine'
  • Comparing models or prompts with numbers instead of anecdotes
  • Catching quality regressions before customers do

What you give it

  • The task your AI feature performs and examples of success and failure
  • Judgement calls on the rubric: what counts as good, what is disqualifying

What you get back

  • A test set built from real cases, labelled and version-controlled
  • Metrics that fit the task — not one generic score hiding three failure modes
  • A repeatable harness with a regression gate: change the prompt or model, see the delta

How it works

  1. Builds the test set from real traffic and real failures, stratified so the hard cases cannot hide in the average.
  2. Designs metrics per failure mode: correctness, faithfulness, format, tone — scored separately.
  3. Uses the right judge per metric: exact checks where possible, rubric-guided model grading where not, with the judge itself spot-checked against human labels.
  4. Runs repeatedly (variance is real) and reports distributions, not single lucky runs.
  5. Installs the harness as a gate: every prompt/model change runs it, deltas reported per metric and per case slice.

Example

You: We want to switch to a cheaper model for our document summariser but we are scared.

Result: A 150-case set from real documents (the long ones, the tables, the scanned mess included), a rubric scoring coverage, faithfulness and length separately, both models run three times each: the cheaper model matches on coverage, wins on length discipline, loses 4 points on faithfulness with tables — decision made with eyes open, and the harness stays as the regression gate.

Limits — please read

  • An evaluation is only as good as its cases; it will keep pushing real, ugly data into the set.
  • Model-graded metrics carry judge bias — it measures and reports that gap rather than pretending it away.
  • Offline scores do not capture everything; it will tell you which claims need an online check with real users.