Agents / Skills

A/B Test Design

Skill

Designs an experiment that can actually answer its question: the hypothesis sharpened, the sample size computed honestly, the metrics and guardrails chosen, and the decision rule written before the data can argue back.

Best for

  • Tests currently designed as 'ship it to half and see'
  • Finding out the test you planned cannot detect the effect you hope for — before running it
  • Making 'the test said so' actually mean something

What you give it

  • The change you want to test, your traffic reality, and the metric you hope to move

What you get back

  • The design: hypothesis with its expected mechanism, primary metric, minimum detectable effect, computed duration at your traffic
  • The guardrails: the metrics that must not degrade, with their stop conditions
  • The decision rule, pre-committed: what result ships it, kills it, or extends it — written before launch, because post-hoc rules always say ship

How it works

  1. Sharpens the hypothesis to a mechanism: what changes, moving which behaviour, visible in which metric — vague hopes design vague tests.
  2. Runs the power math upfront at your real traffic: the minimum detectable effect and duration, because the commonest failure is a test that could never have seen its answer.
  3. Locks the primary metric (one), supports it with mechanism metrics, and fences it with guardrails that have stop conditions.
  4. Pre-commits the decision rule and the analysis plan — peeking rules, segment rules, the ship/kill/extend thresholds — before data exists to be argued with.

Example

You: We want to test the new checkout flow. Design it properly.

Result: The honest math first: at your traffic, detecting the hoped-for 2% conversion lift needs 7 weeks — surfaced BEFORE launch, with the options (test a bolder variant, accept 4 weeks for detecting 3%+, or run on the higher-traffic segment). Chosen: 4 weeks, 3% MDE. The design: hypothesis with mechanism (fewer form fields → fewer abandonments at the address step — so the step-level funnel is instrumented too), primary metric locked (completed checkouts per session entering checkout), guardrails set (average order value, support contacts, refund rate — with stop thresholds), randomisation by account not session (the contamination everyone forgets), and the decision rule signed before launch. Week 4: +3.8% on primary, guardrails clean — shipped, by the rule everyone had already agreed to.

Limits — please read

  • Small traffic means honest constraints: bigger effects only, longer runs, or sequential designs — it says which, with numbers.
  • Randomisation and contamination pitfalls (shared accounts, cross-device) are designed around where your product allows; residual risks are named.
  • A clean test answers its question; what to test next is strategy — the design includes what this test cannot tell you.