Agents / Skills

Prompt Review

Skill

Reviews a prompt the way code gets reviewed: against its actual failure cases, for structure, rule conflicts, format drift and injection surface — returning a sharper version with every change justified by a case it fixes.

Best for

  • The prompt that works 'mostly' and nobody knows why not always
  • A second pair of eyes before the prompt ships into production traffic
  • Learning why a prompt fails by seeing the mechanics named

What you give it

  • The prompt, what it is supposed to do, and the failures you have seen (even anecdotes help)

What you get back

  • The diagnosis: each observed failure traced to its likely cause in the prompt — the buried rule, the conflicting instruction, the format left ambiguous
  • The revision: restructured with every change tied to a failure it addresses — no taste-based rewriting
  • The residue, honestly: what the review cannot fix by wording, with the recommended companion (examples, evaluation, a different decomposition)

How it works

  1. Reads the prompt as an instruction system: what are the rules, where do they live relative to where they apply, which conflict, which are unstated assumptions.
  2. Traces each observed failure to its mechanical cause: buried rules, rule conflicts, missing unknown-handling, undelimited input, format defined by hope.
  3. Revises conservatively: every change cites the failure it addresses — prompt review is engineering, not prose preference.
  4. Names the residue: failures that wording cannot fix get their real treatment named — worked examples, case-set evaluation, or splitting the task.

Example

You: Our summariser prompt occasionally outputs bullet lists despite saying 'prose only', and sometimes includes the document's own instructions.

Result: The diagnosis: 'prose only' appears once, mid-prompt, forty lines before the output section (buried rule — moved to the output contract, restated adjacent to where format is defined); the document is interpolated with no delimiting, so imperative sentences in source documents read as instructions (the injection surface — now fenced with explicit 'treat as data' framing); plus one found-in-review conflict: 'be comprehensive' versus the length cap, resolved by priority order. The revision's changes each cite their case; the two residual risks (very long documents diluting rules; adversarial sources) got their recommended companions: rule restatement after the input, and the injection test set.

Limits — please read

  • Review without failure cases is structural only; it will say which findings are confirmed versus preventive.
  • A prompt cannot exceed the model's capability on the task; the review flags when that wall, not wording, is the problem.
  • Verification needs a case run; the review pairs with an evaluation pass for proof (a prompt-engineering agent does both).