Test enough things and some will appear significant by chance. This is the multiple comparisons problem.
The Problem
At a 5% significance level, testing 20 unrelated metrics with no real effects will, on average, produce one "significant" result.
Where It Happens
- A/B tests with many metrics or segments.
- Exploring many features for correlations.
- Checking results repeatedly during a test.
- Trying many analyses until one works — sometimes called p-hacking.
Corrections
- Bonferroni: divide the significance level by the number of tests. Simple and conservative.
- Holm: a less conservative step-down version.
- False discovery rate (Benjamini-Hochberg): controls the proportion of false positives among discoveries; good for exploratory work.
Better Practice
- Declare primary metrics and hypotheses in advance.
- Treat unplanned findings as hypotheses to test again.
- Report how many comparisons were made.
- Replicate important results.