Best for
- Tests currently designed as 'ship it to half and see'
- Finding out the test you planned cannot detect the effect you hope for — before running it
- Making 'the test said so' actually mean something
What you give it
- The change you want to test, your traffic reality, and the metric you hope to move
What you get back
- The design: hypothesis with its expected mechanism, primary metric, minimum detectable effect, computed duration at your traffic
- The guardrails: the metrics that must not degrade, with their stop conditions
- The decision rule, pre-committed: what result ships it, kills it, or extends it — written before launch, because post-hoc rules always say ship
How it works
- Sharpens the hypothesis to a mechanism: what changes, moving which behaviour, visible in which metric — vague hopes design vague tests.
- Runs the power math upfront at your real traffic: the minimum detectable effect and duration, because the commonest failure is a test that could never have seen its answer.
- Locks the primary metric (one), supports it with mechanism metrics, and fences it with guardrails that have stop conditions.
- Pre-commits the decision rule and the analysis plan — peeking rules, segment rules, the ship/kill/extend thresholds — before data exists to be argued with.
Example
You: We want to test the new checkout flow. Design it properly.
Result: The honest math first: at your traffic, detecting the hoped-for 2% conversion lift needs 7 weeks — surfaced BEFORE launch, with the options (test a bolder variant, accept 4 weeks for detecting 3%+, or run on the higher-traffic segment). Chosen: 4 weeks, 3% MDE. The design: hypothesis with mechanism (fewer form fields → fewer abandonments at the address step — so the step-level funnel is instrumented too), primary metric locked (completed checkouts per session entering checkout), guardrails set (average order value, support contacts, refund rate — with stop thresholds), randomisation by account not session (the contamination everyone forgets), and the decision rule signed before launch. Week 4: +3.8% on primary, guardrails clean — shipped, by the rule everyone had already agreed to.
Limits — please read
- Small traffic means honest constraints: bigger effects only, longer runs, or sequential designs — it says which, with numbers.
- Randomisation and contamination pitfalls (shared accounts, cross-device) are designed around where your product allows; residual risks are named.
- A clean test answers its question; what to test next is strategy — the design includes what this test cannot tell you.