# P-Hacking and Multiple Testing

> Testing many strategy variations guarantees some look good by chance. Learn how p hacking happens, how to adjust for multiple tests and the deflated Sharpe ratio.

Source: https://learn.tradelabsai.com/research/p-hacking-and-multiple-testing/  
Track: Research and Backtesting · Level: Intermediate · Updated: 2026-10-03  
Publisher: TradeLabs AI (https://tradelabsai.com). Education, not financial advice.  
Cite as: TradeLabs Learn, "P-Hacking and Multiple Testing", https://learn.tradelabsai.com/research/p-hacking-and-multiple-testing/

If you test enough trading ideas, some will look profitable purely by chance. This is the multiple testing problem. P hacking is the practice, often unintentional, of trying many analyses, variations or data choices until a result looks statistically significant, then reporting only that result. It is one of the main reasons published anomalies weaken and backtested strategies fail. Accounting for how many things you tried is essential for honest research.

## How many false positives to expect

```
probability of at least one false positive = 1 - (1 - α)^k
```

| Independent tests (k) | Chance of at least one result with p < 0.05 by luck |
|---|---|
| 1 | 5% |
| 10 | 40% |
| 20 | 64% |
| 100 | 99.4% |

With 100 tests of strategies that have no edge, you expect about 5 to look significant.

## Forms of p hacking in trading

| Practice | Example |
|---|---|
| Parameter sweeping | Trying hundreds of indicator settings |
| Rule tinkering | Adding filters until losses disappear |
| Period selection | Choosing start and end dates that look best |
| Market selection | Testing on 30 markets, reporting the 5 that worked |
| Outcome switching | Changing the performance metric to one that looks good |
| Stopping when it looks good | Ending research the moment a backtest succeeds |

## The factor zoo

By the mid 2010s, academic research had reported hundreds of variables that supposedly predict stock returns. Campbell Harvey, Yan Liu and Heqing Zhu (2016) argued that, given how many had been tested, a new factor should have a t statistic above about 3, not the traditional 2. Later replication studies found that many published anomalies fail to replicate or shrink sharply, while others hold up. See [Factor Investing Explained](https://learn.tradelabsai.com/research/factor-investing-explained/).

## Adjusting for multiple tests

| Method | Idea |
|---|---|
| Bonferroni | Require p < α / k; very strict |
| Holm | A stepwise, less strict version of Bonferroni |
| Benjamini Hochberg | Controls the false discovery rate instead of any false positive |
| Higher t thresholds | Use t > 3 for new findings |
| White's Reality Check and Hansen's SPA test | Bootstrap tests for the best of many strategies. See [Bootstrap and Permutation Tests](https://learn.tradelabsai.com/math/bootstrap-and-permutation-tests/) |
| Deflated Sharpe ratio | Adjusts a Sharpe ratio for the number of trials, skew and kurtosis |

## The deflated Sharpe ratio

Bailey and López de Prado (2014) showed that the highest Sharpe ratio among many trials is expected to be positive even if all strategies have zero true Sharpe. The expected maximum grows with the number of trials.

**Example: The best of many random strategies**
Suppose you test 1,000 strategies with no real edge over 5 years of daily data. Because of chance, the best of them is expected to show an annualised Sharpe ratio of around 1.4 in the backtest, and often higher. Seeing a backtest Sharpe of 1.5 therefore means very little if it was the best of 1,000 tries. The deflated Sharpe ratio estimates the probability that a strategy's true Sharpe is above zero after accounting for all those trials. See [Sharpe Ratio](https://learn.tradelabsai.com/portfolio/sharpe-ratio/).

## Honest research practices

1. **Write down hypotheses before testing.**
2. **Log every test,** including failures, with parameters and results.
3. **Count trials** and adjust significance thresholds.
4. **Keep a final holdout** used only once. See [In-Sample vs Out-of-Sample Testing](https://learn.tradelabsai.com/research/out-of-sample-testing/).
5. **Prefer ideas with economic rationale.**
6. **Report all results,** not just the best.
7. **Replicate** on other markets and periods.

## Frequently asked questions

### What is p hacking in trading?

Trying many strategy variations, data choices or analyses until one looks significant, then presenting it as if it were the only test.

### Why is multiple testing a problem?

Because the more ideas you test, the more likely some look profitable by chance, creating false confidence.

### How do I adjust for testing many strategies?

Count all trials and use stricter thresholds, false discovery rate methods, bootstrap tests or the deflated Sharpe ratio, and validate on untouched data.

Next, learn to simulate many possible outcomes in [Monte Carlo Simulation](https://learn.tradelabsai.com/research/monte-carlo-simulation/).

## Continue learning

- Next lesson: [Monte Carlo Simulation](https://learn.tradelabsai.com/research/monte-carlo-simulation/)
- Previous lesson: [Data Leakage](https://learn.tradelabsai.com/research/data-leakage/)
- Related: [Data Leakage](https://learn.tradelabsai.com/research/data-leakage/): Data leakage lets information from test data or the future slip into model training. Learn common leaks in trading and machine learning and how to prevent them.
- Related: [Statistical Significance in Trading](https://learn.tradelabsai.com/math/statistical-significance/): Statistical significance helps judge whether trading results reflect a real edge or luck. Learn the t statistic rule of thumb, sample size and multiple testing.
- Related: [Overfitting and Curve Fitting](https://learn.tradelabsai.com/research/overfitting-and-curve-fitting/): Overfitting means a strategy fits noise instead of a real pattern. Learn the warning signs, why it happens, how to measure it and practical ways to avoid it.
- Related: [Hypothesis Testing and P-Values](https://learn.tradelabsai.com/math/hypothesis-testing-and-p-values/): Hypothesis tests check whether results are likely due to chance. Learn null hypotheses, test statistics and p values, a strategy test and how p values mislead.
- Related: [Statistical Power and Type I and II Errors](https://learn.tradelabsai.com/math/type-i-and-ii-errors/): Type I errors are false positives; type II errors are missed real effects. Learn how they apply to strategy testing, the trade off between them and power.
- Related: [Backtest Reproducibility](https://learn.tradelabsai.com/research/backtest-reproducibility/): A reproducible backtest gives the same results every time from the same code and data. Learn version control, data snapshots, research logs and good habits.
