P-Hacking and Multiple Testing
Testing many strategy variations guarantees some look good by chance. Learn how p hacking happens, how to adjust for multiple tests and the deflated Sharpe ratio.
If you test enough trading ideas, some will look profitable purely by chance. This is the multiple testing problem. P hacking is the practice, often unintentional, of trying many analyses, variations or data choices until a result looks statistically significant, then reporting only that result. It is one of the main reasons published anomalies weaken and backtested strategies fail. Accounting for how many things you tried is essential for honest research.
How many false positives to expect#
probability of at least one false positive = 1 - (1 - α)^k
| Independent tests (k) | Chance of at least one result with p < 0.05 by luck |
|---|---|
| 1 | 5% |
| 10 | 40% |
| 20 | 64% |
| 100 | 99.4% |
With 100 tests of strategies that have no edge, you expect about 5 to look significant.
Forms of p hacking in trading#
| Practice | Example |
|---|---|
| Parameter sweeping | Trying hundreds of indicator settings |
| Rule tinkering | Adding filters until losses disappear |
| Period selection | Choosing start and end dates that look best |
| Market selection | Testing on 30 markets, reporting the 5 that worked |
| Outcome switching | Changing the performance metric to one that looks good |
| Stopping when it looks good | Ending research the moment a backtest succeeds |
The factor zoo#
By the mid 2010s, academic research had reported hundreds of variables that supposedly predict stock returns. Campbell Harvey, Yan Liu and Heqing Zhu (2016) argued that, given how many had been tested, a new factor should have a t statistic above about 3, not the traditional 2. Later replication studies found that many published anomalies fail to replicate or shrink sharply, while others hold up. See Factor Investing Explained.
Adjusting for multiple tests#
| Method | Idea |
|---|---|
| Bonferroni | Require p < α / k; very strict |
| Holm | A stepwise, less strict version of Bonferroni |
| Benjamini Hochberg | Controls the false discovery rate instead of any false positive |
| Higher t thresholds | Use t > 3 for new findings |
| White's Reality Check and Hansen's SPA test | Bootstrap tests for the best of many strategies. See Bootstrap and Permutation Tests |
| Deflated Sharpe ratio | Adjusts a Sharpe ratio for the number of trials, skew and kurtosis |
The deflated Sharpe ratio#
Bailey and López de Prado (2014) showed that the highest Sharpe ratio among many trials is expected to be positive even if all strategies have zero true Sharpe. The expected maximum grows with the number of trials.
Honest research practices#
- Write down hypotheses before testing.
- Log every test, including failures, with parameters and results.
- Count trials and adjust significance thresholds.
- Keep a final holdout used only once. See In-Sample vs Out-of-Sample Testing.
- Prefer ideas with economic rationale.
- Report all results, not just the best.
- Replicate on other markets and periods.
Frequently asked questions#
What is p hacking in trading?#
Trying many strategy variations, data choices or analyses until one looks significant, then presenting it as if it were the only test.
Why is multiple testing a problem?#
Because the more ideas you test, the more likely some look profitable by chance, creating false confidence.
How do I adjust for testing many strategies?#
Count all trials and use stricter thresholds, false discovery rate methods, bootstrap tests or the deflated Sharpe ratio, and validate on untouched data.
Next, learn to simulate many possible outcomes in Monte Carlo Simulation.
3 quick questions on this lesson. Get them all right to finish it.
Turn on JavaScript to take the quiz.
Mentioned in
- The Trading Research ProcessResearch and Backtesting
- Why Strategies FailResearch and Backtesting
- Backtesting MethodologyResearch and Backtesting
- In-Sample vs Out-of-Sample TestingResearch and Backtesting
- Survivorship and Selection BiasResearch and Backtesting
- Data LeakageResearch and Backtesting