Best for
- The 'should we fine-tune?' meeting that recurs quarterly without a framework
- Budgeting an AI feature before committing to its most expensive version
- Diagnosing why the current approach underperforms before doubling down on it
What you give it
- The capability needed, examples of current failures (if any), and your constraints — data, budget, latency, update frequency
What you get back
- The diagnosis: what the capability actually requires — knowledge, behaviour, format, or freshness — because each points at a different tool
- The recommendation with its reasoning: the path (usually staged), what it costs, what it cannot fix, and the exit criteria for escalating to the next path
- The decision record: why this path, what was rejected and why — so the quarterly meeting stops recurring
How it works
- Decomposes the requirement by what is actually missing: knowledge the model lacks, behaviour it will not consistently produce, format discipline, or freshness — the taxonomy that maps to tools.
- Maps honestly: fresh or proprietary knowledge wants retrieval; consistent behaviour and format want prompting with examples first and fine-tuning only past proven prompt limits; freshness rules out baking knowledge into weights.
- Prices each path completely: not just build cost but the maintenance tail — the stale fine-tune, the retrieval index's freshness pipeline, the prompt's evaluation harness.
- Stages the recommendation with exit criteria: the cheap path first with a measurable bar, escalation defined by evidence, so the decision executes itself instead of re-arguing.
Example
You: Our product-support answers are mediocre. The team wants to fine-tune. Decide.
Result: The diagnosis said otherwise: the failures decomposed into missing knowledge (product details the model cannot know — a retrieval problem; fine-tuning bakes in knowledge that goes stale with every release) and inconsistent tone (a behaviour problem — but one that better prompting with worked examples fixes at a fraction of the cost). The staged recommendation: retrieval over the product docs first (two weeks, measurable), prompt engineering for tone with an eval set (one week), and fine-tuning held as the explicit escalation IF the eval still shows behaviour gaps prompting cannot close — with its real costs stated (data curation, retraining per model generation, the evaluation burden). Outcome: the first two stages hit the quality bar; the fine-tune budget was never spent; the decision record ended the quarterly debate.
Limits — please read
- The diagnosis needs real failure examples; 'make it better' without cases gets the case-collection step first.
- Costs are estimated against your stated constraints; vendor pricing shifts, and the record says which numbers to refresh.
- Genuine fine-tune territory exists (deep style, narrow domains at scale, latency-critical distillation); the framework recognises it rather than reflexively avoiding it.