Skills
A skill is a reusable procedure your AI tool follows when you ask — an impact analysis, a change record, a prompt rebuild.
- SkillUSD 2
Metric Definition
Defines a business metric so precisely that two people compute the same number: numerator, denominator, filters, time window, edge cases — settled once, written down, and reconciled against the numbers already in circulation.
Best for: The meeting where three dashboards show three 'active users' numbers
- SkillUSD 2
Dashboard Spec
Specifies a dashboard before anyone builds it: the questions it answers, for whom, with which metrics at which grain — so what gets built is an instrument for decisions, not another wall of charts nobody reads twice.
Best for: The dashboard request that is actually a vague anxiety
- SkillUSD 2
A/B Test Design
Designs an experiment that can actually answer its question: the hypothesis sharpened, the sample size computed honestly, the metrics and guardrails chosen, and the decision rule written before the data can argue back.
Best for: Tests currently designed as 'ship it to half and see'
- SkillUSD 2
Prompt Review
Reviews a prompt the way code gets reviewed: against its actual failure cases, for structure, rule conflicts, format drift and injection surface — returning a sharper version with every change justified by a case it fixes.
Best for: The prompt that works 'mostly' and nobody knows why not always
- SkillUSD 2
Prompt Template Builder
Turns a task you keep prompting ad hoc into a reusable template: the stable instruction frame engineered once, the variable slots defined with their rules, and the filled examples that show the next user exactly how to use it.
Best for: The task you re-explain to the model every single time
- SkillUSD 2
Eval Dataset Builder
Builds the test set your AI feature is judged by: real cases collected and labelled, the hard and ugly ones deliberately included, stratified so the average cannot hide the failures — the ground truth that makes quality measurable.
Best for: AI features evaluated today on five cherry-picked examples
- SkillUSD 2
LLM Judge Rubric
Builds a model-graded evaluation you can actually trust: the rubric decomposed into judgeable criteria, the judge prompt engineered against its known biases, and the judge itself validated against human labels before its scores mean anything.
Best for: Grading a thousand outputs where humans can grade fifty
- SkillUSD 2
RAG Chunking Strategy
Chooses how your documents get split for retrieval — by structure, size and overlap fitted to what your documents actually are — and proves the choice against real queries, because chunking is where most RAG quality is silently won or lost.
Best for: RAG systems built on the default 500-tokens-and-hope
- SkillUSD 2
Retrieval Quality Check
Measures whether your RAG system's retrieval actually finds the right passages — separately from whether the answers read well — because a fluent answer from the wrong source is the failure mode that erodes trust silently.
Best for: RAG debugging that currently starts (and ends) at the final answer
- SkillUSD 2
Grounding Check
Verifies that your AI feature's answers actually come from their sources: claims traced to the retrieved or provided material, inventions caught and classified, and the honest-refusal behaviour tested — the fluency-versus-faithfulness audit.
Best for: RAG and summarisation features that sound right more reliably than they are right
Also see the Agents.