Agents / Skills

Prompt vs RAG vs Fine-tune

Skill

Decides how an AI capability should be built — better prompting, retrieval over your data, fine-tuning, or a combination — from the actual requirement, with the costs and failure modes of each path stated before money gets spent on the wrong one.

Best for

  • The 'should we fine-tune?' meeting that recurs quarterly without a framework
  • Budgeting an AI feature before committing to its most expensive version
  • Diagnosing why the current approach underperforms before doubling down on it

What you give it

  • The capability needed, examples of current failures (if any), and your constraints — data, budget, latency, update frequency

What you get back

  • The diagnosis: what the capability actually requires — knowledge, behaviour, format, or freshness — because each points at a different tool
  • The recommendation with its reasoning: the path (usually staged), what it costs, what it cannot fix, and the exit criteria for escalating to the next path
  • The decision record: why this path, what was rejected and why — so the quarterly meeting stops recurring

How it works

  1. Decomposes the requirement by what is actually missing: knowledge the model lacks, behaviour it will not consistently produce, format discipline, or freshness — the taxonomy that maps to tools.
  2. Maps honestly: fresh or proprietary knowledge wants retrieval; consistent behaviour and format want prompting with examples first and fine-tuning only past proven prompt limits; freshness rules out baking knowledge into weights.
  3. Prices each path completely: not just build cost but the maintenance tail — the stale fine-tune, the retrieval index's freshness pipeline, the prompt's evaluation harness.
  4. Stages the recommendation with exit criteria: the cheap path first with a measurable bar, escalation defined by evidence, so the decision executes itself instead of re-arguing.

Example

You: Our product-support answers are mediocre. The team wants to fine-tune. Decide.

Result: The diagnosis said otherwise: the failures decomposed into missing knowledge (product details the model cannot know — a retrieval problem; fine-tuning bakes in knowledge that goes stale with every release) and inconsistent tone (a behaviour problem — but one that better prompting with worked examples fixes at a fraction of the cost). The staged recommendation: retrieval over the product docs first (two weeks, measurable), prompt engineering for tone with an eval set (one week), and fine-tuning held as the explicit escalation IF the eval still shows behaviour gaps prompting cannot close — with its real costs stated (data curation, retraining per model generation, the evaluation burden). Outcome: the first two stages hit the quality bar; the fine-tune budget was never spent; the decision record ended the quarterly debate.

Limits — please read

  • The diagnosis needs real failure examples; 'make it better' without cases gets the case-collection step first.
  • Costs are estimated against your stated constraints; vendor pricing shifts, and the record says which numbers to refresh.
  • Genuine fine-tune territory exists (deep style, narrow domains at scale, latency-critical distillation); the framework recognises it rather than reflexively avoiding it.