TradeLabs AILearn

P-Hacking and Multiple Testing

Testing many strategy variations guarantees some look good by chance. Learn how p hacking happens, how to adjust for multiple tests and the deflated Sharpe ratio.

Intermediate3 min readUpdated 3 Oct 2026
Markdown
Read firstData Leakage
Lesson 12 of 38

If you test enough trading ideas, some will look profitable purely by chance. This is the multiple testing problem. P hacking is the practice, often unintentional, of trying many analyses, variations or data choices until a result looks statistically significant, then reporting only that result. It is one of the main reasons published anomalies weaken and backtested strategies fail. Accounting for how many things you tried is essential for honest research.

How many false positives to expect#

probability of at least one false positive = 1 - (1 - α)^k
Independent tests (k)Chance of at least one result with p < 0.05 by luck
15%
1040%
2064%
10099.4%

With 100 tests of strategies that have no edge, you expect about 5 to look significant.

Forms of p hacking in trading#

PracticeExample
Parameter sweepingTrying hundreds of indicator settings
Rule tinkeringAdding filters until losses disappear
Period selectionChoosing start and end dates that look best
Market selectionTesting on 30 markets, reporting the 5 that worked
Outcome switchingChanging the performance metric to one that looks good
Stopping when it looks goodEnding research the moment a backtest succeeds

The factor zoo#

By the mid 2010s, academic research had reported hundreds of variables that supposedly predict stock returns. Campbell Harvey, Yan Liu and Heqing Zhu (2016) argued that, given how many had been tested, a new factor should have a t statistic above about 3, not the traditional 2. Later replication studies found that many published anomalies fail to replicate or shrink sharply, while others hold up. See Factor Investing Explained.

Adjusting for multiple tests#

MethodIdea
BonferroniRequire p < α / k; very strict
HolmA stepwise, less strict version of Bonferroni
Benjamini HochbergControls the false discovery rate instead of any false positive
Higher t thresholdsUse t > 3 for new findings
White's Reality Check and Hansen's SPA testBootstrap tests for the best of many strategies. See Bootstrap and Permutation Tests
Deflated Sharpe ratioAdjusts a Sharpe ratio for the number of trials, skew and kurtosis

The deflated Sharpe ratio#

Bailey and López de Prado (2014) showed that the highest Sharpe ratio among many trials is expected to be positive even if all strategies have zero true Sharpe. The expected maximum grows with the number of trials.

Honest research practices#

  1. Write down hypotheses before testing.
  2. Log every test, including failures, with parameters and results.
  3. Count trials and adjust significance thresholds.
  4. Keep a final holdout used only once. See In-Sample vs Out-of-Sample Testing.
  5. Prefer ideas with economic rationale.
  6. Report all results, not just the best.
  7. Replicate on other markets and periods.

Frequently asked questions#

What is p hacking in trading?#

Trying many strategy variations, data choices or analyses until one looks significant, then presenting it as if it were the only test.

Why is multiple testing a problem?#

Because the more ideas you test, the more likely some look profitable by chance, creating false confidence.

How do I adjust for testing many strategies?#

Count all trials and use stricter thresholds, false discovery rate methods, bootstrap tests or the deflated Sharpe ratio, and validate on untouched data.

Next, learn to simulate many possible outcomes in Monte Carlo Simulation.

Check your understanding

3 quick questions on this lesson. Get them all right to finish it.

Turn on JavaScript to take the quiz.

Finished this lesson?Sign in to save your progress across devices.
Next lessonMonte Carlo SimulationMonte Carlo simulation generates thousands of possible outcomes to show the range of results. Learn trade resampling, drawdown estimates and the limits.

Mentioned in