Agents

Skills

A skill is a reusable procedure your AI tool follows when you ask — an impact analysis, a change record, a prompt rebuild.

  • SkillUSD 2

    Metric Definition

    Defines a business metric so precisely that two people compute the same number: numerator, denominator, filters, time window, edge cases — settled once, written down, and reconciled against the numbers already in circulation.

    Best for: The meeting where three dashboards show three 'active users' numbers

  • SkillUSD 2

    Dashboard Spec

    Specifies a dashboard before anyone builds it: the questions it answers, for whom, with which metrics at which grain — so what gets built is an instrument for decisions, not another wall of charts nobody reads twice.

    Best for: The dashboard request that is actually a vague anxiety

  • SkillUSD 2

    A/B Test Design

    Designs an experiment that can actually answer its question: the hypothesis sharpened, the sample size computed honestly, the metrics and guardrails chosen, and the decision rule written before the data can argue back.

    Best for: Tests currently designed as 'ship it to half and see'

  • SkillUSD 2

    Prompt Review

    Reviews a prompt the way code gets reviewed: against its actual failure cases, for structure, rule conflicts, format drift and injection surface — returning a sharper version with every change justified by a case it fixes.

    Best for: The prompt that works 'mostly' and nobody knows why not always

  • SkillUSD 2

    Prompt Template Builder

    Turns a task you keep prompting ad hoc into a reusable template: the stable instruction frame engineered once, the variable slots defined with their rules, and the filled examples that show the next user exactly how to use it.

    Best for: The task you re-explain to the model every single time

  • SkillUSD 2

    Eval Dataset Builder

    Builds the test set your AI feature is judged by: real cases collected and labelled, the hard and ugly ones deliberately included, stratified so the average cannot hide the failures — the ground truth that makes quality measurable.

    Best for: AI features evaluated today on five cherry-picked examples

  • SkillUSD 2

    LLM Judge Rubric

    Builds a model-graded evaluation you can actually trust: the rubric decomposed into judgeable criteria, the judge prompt engineered against its known biases, and the judge itself validated against human labels before its scores mean anything.

    Best for: Grading a thousand outputs where humans can grade fifty

  • SkillUSD 2

    RAG Chunking Strategy

    Chooses how your documents get split for retrieval — by structure, size and overlap fitted to what your documents actually are — and proves the choice against real queries, because chunking is where most RAG quality is silently won or lost.

    Best for: RAG systems built on the default 500-tokens-and-hope

  • SkillUSD 2

    Retrieval Quality Check

    Measures whether your RAG system's retrieval actually finds the right passages — separately from whether the answers read well — because a fluent answer from the wrong source is the failure mode that erodes trust silently.

    Best for: RAG debugging that currently starts (and ends) at the final answer

  • SkillUSD 2

    Grounding Check

    Verifies that your AI feature's answers actually come from their sources: claims traced to the retrieved or provided material, inventions caught and classified, and the honest-refusal behaviour tested — the fluency-versus-faithfulness audit.

    Best for: RAG and summarisation features that sound right more reliably than they are right

Also see the Agents.