# Hypothesis Testing and P-Values

> Hypothesis tests check whether results are likely due to chance. Learn null hypotheses, test statistics and p values, a strategy test and how p values mislead.

Source: https://learn.tradelabsai.com/math/hypothesis-testing-and-p-values/  
Track: Math and Statistics · Level: Intermediate · Updated: 2026-10-03  
Publisher: TradeLabs AI (https://tradelabsai.com). Education, not financial advice.  
Cite as: TradeLabs Learn, "Hypothesis Testing and P-Values", https://learn.tradelabsai.com/math/hypothesis-testing-and-p-values/

When a backtest shows profits, the key question is whether the strategy has a real edge or just got lucky. Hypothesis testing is a formal way to answer it. You assume the boring explanation (no edge) and ask how surprising your results would be if that were true. The p value measures that surprise. Hypothesis tests are widely used in trading research, but they are also widely misunderstood and misused, especially when many strategies are tested.

## The steps

1. **State the null hypothesis (H0):** the default, such as "the strategy's true average return is zero".
2. **State the alternative (H1):** such as "the true average return is greater than zero".
3. **Choose a significance level (α),** often 5%.
4. **Calculate a test statistic** from the data.
5. **Find the p value:** the probability of results at least this extreme if H0 were true.
6. **Decide:** if p < α, reject H0; otherwise, do not reject it.

## The t test for average returns

```
t = (sample mean - hypothesised mean) / (standard deviation / √n)
```

**Example: Testing a strategy**
A strategy has 144 trades with an average return of +0.30% and a standard deviation of 1.8%.

t = (0.30 minus 0) / (1.8 / √144) = 0.30 / 0.15 = 2.0.

For a one sided test with 143 degrees of freedom, the p value is about 0.024. At α = 5%, we reject the null hypothesis: if the strategy truly had zero edge, results this good would occur only about 2.4% of the time by chance.

But if this was the best of 20 strategies tested, the conclusion changes completely. See [P-Hacking and Multiple Testing](https://learn.tradelabsai.com/research/p-hacking-and-multiple-testing/).

## What a p value is and is not

| A p value is | A p value is not |
|---|---|
| The probability of seeing results at least this extreme if the null is true | The probability that the null hypothesis is true |
| A measure of how surprising the data are under "no effect" | The probability that the strategy will work in the future |
| Dependent on sample size | A measure of how large or valuable the effect is |

The American Statistical Association issued a statement in 2016 warning against these common misinterpretations and against treating p < 0.05 as a magic threshold.

## Statistical vs practical significance

With enough data, tiny effects become statistically significant. A strategy averaging 0.01% per trade over 1 million trades may be "significant" but useless after costs. Always ask whether the effect is large enough to matter. See [Statistical Significance in Trading](https://learn.tradelabsai.com/math/statistical-significance/).

## Errors

| | H0 true (no edge) | H0 false (real edge) |
|---|---|---|
| Reject H0 | Type I error (false positive) | Correct |
| Do not reject H0 | Correct | Type II error (false negative) |

See [Statistical Power and Type I and II Errors](https://learn.tradelabsai.com/math/type-i-and-ii-errors/).

## Multiple testing

If you test 100 strategies with no edge at α = 5%, about 5 will look significant by chance. Researchers such as Campbell Harvey, Yan Liu and Heqing Zhu (2016) argued that because so many factors have been tested, new trading factors should clear a much higher bar, such as a t statistic above 3 rather than 2. See [P-Hacking and Multiple Testing](https://learn.tradelabsai.com/research/p-hacking-and-multiple-testing/).

## Assumptions to check

- **Independence:** correlated returns overstate significance. See [Autocorrelation and Partial Autocorrelation](https://learn.tradelabsai.com/math/autocorrelation/).
- **Distribution:** t tests assume roughly normal means; heavy tails can distort results. See [Fat Tails](https://learn.tradelabsai.com/math/fat-tails/).
- **Stable process:** a changing market undermines tests on historical data. See [Structural Breaks and Regime Changes](https://learn.tradelabsai.com/math/regime-changes/).

When assumptions fail, bootstrap and permutation tests are alternatives. See [Bootstrap and Permutation Tests](https://learn.tradelabsai.com/math/bootstrap-and-permutation-tests/).

## Frequently asked questions

### What is a p value?

The probability of getting results at least as extreme as those observed, assuming the null hypothesis, such as "no edge", is true.

### What does p < 0.05 mean?

That results this extreme would occur less than 5% of the time if there were no real effect; it does not prove the strategy works.

### Why can hypothesis tests mislead traders?

Because testing many strategies produces false positives, assumptions such as independence may fail, and statistical significance does not guarantee practical value.

Next, learn the two ways tests can be wrong in [Statistical Power and Type I and II Errors](https://learn.tradelabsai.com/math/type-i-and-ii-errors/).

## Continue learning

- Next lesson: [Statistical Power and Type I and II Errors](https://learn.tradelabsai.com/math/type-i-and-ii-errors/)
- Previous lesson: [Confidence Intervals](https://learn.tradelabsai.com/math/confidence-intervals/)
- Related: [Confidence Intervals](https://learn.tradelabsai.com/math/confidence-intervals/): A confidence interval gives a range of plausible values for a statistic. Learn how to calculate them for returns and win rates and how to read them in backtests.
- Related: [Statistical Significance in Trading](https://learn.tradelabsai.com/math/statistical-significance/): Statistical significance helps judge whether trading results reflect a real edge or luck. Learn the t statistic rule of thumb, sample size and multiple testing.
- Related: [Statistical Power and Type I and II Errors](https://learn.tradelabsai.com/math/type-i-and-ii-errors/): Type I errors are false positives; type II errors are missed real effects. Learn how they apply to strategy testing, the trade off between them and power.
- Related: [P-Hacking and Multiple Testing](https://learn.tradelabsai.com/research/p-hacking-and-multiple-testing/): Testing many strategy variations guarantees some look good by chance. Learn how p hacking happens, how to adjust for multiple tests and the deflated Sharpe ratio.
- Related: [Bootstrap and Permutation Tests](https://learn.tradelabsai.com/math/bootstrap-and-permutation-tests/): Bootstrap and permutation tests use resampling to measure uncertainty and test significance without strict assumptions. Learn how they work, examples and pitfalls.
