Best for
- The prompt that works 'mostly' and nobody knows why not always
- A second pair of eyes before the prompt ships into production traffic
- Learning why a prompt fails by seeing the mechanics named
What you give it
- The prompt, what it is supposed to do, and the failures you have seen (even anecdotes help)
What you get back
- The diagnosis: each observed failure traced to its likely cause in the prompt — the buried rule, the conflicting instruction, the format left ambiguous
- The revision: restructured with every change tied to a failure it addresses — no taste-based rewriting
- The residue, honestly: what the review cannot fix by wording, with the recommended companion (examples, evaluation, a different decomposition)
How it works
- Reads the prompt as an instruction system: what are the rules, where do they live relative to where they apply, which conflict, which are unstated assumptions.
- Traces each observed failure to its mechanical cause: buried rules, rule conflicts, missing unknown-handling, undelimited input, format defined by hope.
- Revises conservatively: every change cites the failure it addresses — prompt review is engineering, not prose preference.
- Names the residue: failures that wording cannot fix get their real treatment named — worked examples, case-set evaluation, or splitting the task.
Example
You: Our summariser prompt occasionally outputs bullet lists despite saying 'prose only', and sometimes includes the document's own instructions.
Result: The diagnosis: 'prose only' appears once, mid-prompt, forty lines before the output section (buried rule — moved to the output contract, restated adjacent to where format is defined); the document is interpolated with no delimiting, so imperative sentences in source documents read as instructions (the injection surface — now fenced with explicit 'treat as data' framing); plus one found-in-review conflict: 'be comprehensive' versus the length cap, resolved by priority order. The revision's changes each cite their case; the two residual risks (very long documents diluting rules; adversarial sources) got their recommended companions: rule restatement after the input, and the injection test set.
Limits — please read
- Review without failure cases is structural only; it will say which findings are confirmed versus preventive.
- A prompt cannot exceed the model's capability on the task; the review flags when that wall, not wording, is the problem.
- Verification needs a case run; the review pairs with an evaluation pass for proof (a prompt-engineering agent does both).