All UX laws

Statistical Power

Test duration depends on traffic, baseline rate, and minimum detectable effect.

How to apply it

Use a sample-size calculator before launching. Stopping early inflates false positives.

Statistical Power in practice: a concrete example

The test you stopped reading too early

An A/B test shows the variant up 12% after two days and it is tempting to ship it. But with low traffic and a small baseline rate, that early lead is mostly noise, and calling it now inflates the chance of a false win. Statistical power is about running long enough (traffic, baseline rate, and the effect you want to detect together) to tell a real difference from randomness. Put the numbers into a sample-size calculator before you launch, then wait for it.

Common Statistical Power mistakes

  • Peeking at the results and stopping the moment the line looks good, which sharply inflates false positives.
  • Running a test on traffic too thin to ever detect the effect size you care about, so it is inconclusive by design.
  • Ignoring the minimum detectable effect: a test powered to catch a 20% lift will call a real 3% lift a tie.

When Statistical Power doesn't apply

Formal power analysis is for decisions that hinge on a small, measurable difference. For a large, obvious change, or a product with too little traffic to ever reach significance, waiting months for a p-value is the wrong tool, and a qualitative read or a clear product judgment serves better. Rigour is worth it in proportion to the cost of being wrong.

Practise it on a real challenge

CTA Copy — 'Get started' vs 'Try free for 14 days'

PM wants to test 'Get started' vs 'Try free for 14 days'.

More UX laws

  • 60-30-10 Rule
  • 8-pt Spacing Scale
  • Accessible contrast
  • Aesthetic-Usability Effect
  • Alignment Principle
  • Calibrated Trust
UX QuestUX LawsPricingAboutContactPopular UX laws:Hick's LawFitts's LawDoherty ThresholdMiller's LawLaw of ProximityAesthetic-Usability EffectVisual hierarchyAccessible contrastAll UX laws