Best for
- Products with an LLM feature that works 'usually'
- Turning a prompt that grew by patching into something maintainable
- Getting consistent output format and tone at scale
What you give it
- What the prompt must do, with examples of good and bad output
- The failure cases you have seen — they become the test set
What you get back
- A restructured prompt: role, rules, output contract, examples — each part there for a reason
- An evaluation run against your real cases, before/after, with the score honest
- Versioned prompts with the change log, so regressions can be found and rolled back
How it works
- Collects real cases first — the prompt is engineered against evidence, not taste.
- Structures the prompt: role, hard rules, output contract, worked examples, the input clearly delimited.
- Tests changes against the case set and reports the score movement honestly.
- Hardens against the classics: instruction-ignoring on long inputs, format drift, hallucinated specifics, prompt injection via user content.
- Versions everything with reasons, so 'it got worse' is answerable with a diff.
Example
You: Our support-reply drafter ignores the refund policy about twice a day and sometimes invents order numbers.
Result: Diagnosis: policy buried mid-prompt, no instruction about unknown facts. Rebuilt with the policy as numbered rules, an explicit 'never state an order number not present in the input' rule, and three worked examples. Evaluated on 60 logged cases: policy violations 9 to 0, invented facts 7 to 1 — the remaining case documented with its workaround.
Limits — please read
- A prompt cannot fix a task the model fundamentally cannot do; it will say when you have hit that wall.
- Scores are only as honest as the case set; it will push you to include the ugly cases.
- Perfect consistency is not available from a probabilistic system — it reduces variance and tells you what remains.