Hypothesis Testing and P-Values
Hypothesis tests check whether results are likely due to chance. Learn null hypotheses, test statistics and p values, a strategy test and how p values mislead.
When a backtest shows profits, the key question is whether the strategy has a real edge or just got lucky. Hypothesis testing is a formal way to answer it. You assume the boring explanation (no edge) and ask how surprising your results would be if that were true. The p value measures that surprise. Hypothesis tests are widely used in trading research, but they are also widely misunderstood and misused, especially when many strategies are tested.
The steps#
- State the null hypothesis (H0): the default, such as "the strategy's true average return is zero".
- State the alternative (H1): such as "the true average return is greater than zero".
- Choose a significance level (α), often 5%.
- Calculate a test statistic from the data.
- Find the p value: the probability of results at least this extreme if H0 were true.
- Decide: if p < α, reject H0; otherwise, do not reject it.
The t test for average returns#
t = (sample mean - hypothesised mean) / (standard deviation / √n)
What a p value is and is not#
| A p value is | A p value is not |
|---|---|
| The probability of seeing results at least this extreme if the null is true | The probability that the null hypothesis is true |
| A measure of how surprising the data are under "no effect" | The probability that the strategy will work in the future |
| Dependent on sample size | A measure of how large or valuable the effect is |
The American Statistical Association issued a statement in 2016 warning against these common misinterpretations and against treating p < 0.05 as a magic threshold.
Statistical vs practical significance#
With enough data, tiny effects become statistically significant. A strategy averaging 0.01% per trade over 1 million trades may be "significant" but useless after costs. Always ask whether the effect is large enough to matter. See Statistical Significance in Trading.
Errors#
| H0 true (no edge) | H0 false (real edge) | |
|---|---|---|
| Reject H0 | Type I error (false positive) | Correct |
| Do not reject H0 | Correct | Type II error (false negative) |
See Statistical Power and Type I and II Errors.
Multiple testing#
If you test 100 strategies with no edge at α = 5%, about 5 will look significant by chance. Researchers such as Campbell Harvey, Yan Liu and Heqing Zhu (2016) argued that because so many factors have been tested, new trading factors should clear a much higher bar, such as a t statistic above 3 rather than 2. See P-Hacking and Multiple Testing.
Assumptions to check#
- Independence: correlated returns overstate significance. See Autocorrelation and Partial Autocorrelation.
- Distribution: t tests assume roughly normal means; heavy tails can distort results. See Fat Tails.
- Stable process: a changing market undermines tests on historical data. See Structural Breaks and Regime Changes.
When assumptions fail, bootstrap and permutation tests are alternatives. See Bootstrap and Permutation Tests.
Frequently asked questions#
What is a p value?#
The probability of getting results at least as extreme as those observed, assuming the null hypothesis, such as "no edge", is true.
What does p < 0.05 mean?#
That results this extreme would occur less than 5% of the time if there were no real effect; it does not prove the strategy works.
Why can hypothesis tests mislead traders?#
Because testing many strategies produces false positives, assumptions such as independence may fail, and statistical significance does not guarantee practical value.
Next, learn the two ways tests can be wrong in Statistical Power and Type I and II Errors.
3 quick questions on this lesson. Get them all right to finish it.
Turn on JavaScript to take the quiz.
Mentioned in
- Central Limit TheoremMath and Statistics
- Student's t-DistributionMath and Statistics
- Quant Trading Learning PathStart Here
- Reading Academic PapersStart Here
- Quant Trader and Quant ResearcherThe Trading Industry