Best for
- Releases too risky for all-at-once
- Teams whose codebase is an archaeology of dead flags
- Making 'we can always turn it off' actually true
What you give it
- The feature, its risk profile, and the metrics that would show trouble
What you get back
- A staged rollout plan: cohorts, percentages, dwell times, and the gate metrics per stage
- Kill-switch semantics defined and tested: what off means for in-flight users and their data
- The removal plan: when the flag dies, who removes it, what reminds them
How it works
- Sizes the staging to the risk: money paths get shadow comparison; cosmetic changes get two stages, not five.
- Defines gates as metrics with thresholds, not vibes — each stage needs its evidence before widening.
- Specifies kill-switch semantics precisely: what off means for users mid-flow and data already written.
- Schedules the flag's removal at plan time, because flags planned without funerals become permanent residents.
Example
You: We are rolling out the new pricing calculation. If it is wrong, we lose money or trust. Plan it.
Result: Five stages: internal accounts, 1% of traffic with old-vs-new shadow comparison (mismatches logged, zero required to proceed), 10% with the error-rate and support-ticket gates, 50%, 100%. Kill-switch tested: off reverts to old calculation, in-flight checkouts complete on whichever version priced their cart. Flag removal scheduled two weeks after 100%, with the cleanup ticket pre-created.
Limits — please read
- It plans within your flag tooling; introducing a flag system is a separate decision.
- Shadow comparisons need the old and new paths to coexist — where they cannot, it says so and adjusts the design.
- Gate metrics need instrumentation that exists; missing metrics are named as prerequisites.