Statistical Significance in Trading
Statistical significance helps judge whether trading results reflect a real edge or luck. Learn the t statistic rule of thumb, sample size and multiple testing.
A strategy that made money in a backtest or a few months of live trading may have a real edge, or it may have been lucky. Statistical significance is a way of judging how unlikely the results would be if there were no edge at all. It is not proof, and it can be badly misused, but it provides a disciplined first filter. This lesson brings together the ideas from earlier lessons into practical rules for judging trading results.
The t statistic as a quick guide#
For average returns:
t = mean return / (standard deviation / √n)
For an annualised Sharpe ratio measured over T years, a convenient approximation is:
t ≈ Sharpe ratio × √T
| t statistic | Rough interpretation (single test, no data mining) |
|---|---|
| Below 1 | Indistinguishable from noise |
| About 2 | Conventionally "significant" at 5% |
| Above 3 | Strong evidence; recommended for new factors after many have been tested |
Sample size matters more than you think#
Statistical significance depends heavily on the number of independent observations. Strategies that trade rarely, such as a few times a year, may never produce enough trades to be statistically convincing within a reasonable time. High frequency strategies accumulate observations quickly, which partly explains why their edges can be measured more precisely. See Sampling and Standard Error.
Multiple testing changes everything#
If you test many ideas, some will look significant by chance. With 20 independent tests of strategies with no edge, the chance that at least one shows p < 0.05 is 1 minus 0.95^20 ≈ 64%.
Ways to adjust:
| Method | Approach |
|---|---|
| Bonferroni correction | Divide α by the number of tests; strict |
| Holm and Benjamini Hochberg | Less strict; control the false discovery rate |
| Higher t thresholds | Such as t > 3, as suggested by Harvey, Liu and Zhu (2016) |
| Deflated Sharpe ratio | Adjusts the Sharpe ratio for the number of trials and non normal returns (Bailey and López de Prado) |
| Out of sample testing | Test on data not used to develop the strategy. See In-Sample vs Out-of-Sample Testing |
See P-Hacking and Multiple Testing.
Significance is not importance#
- Statistical significance says an effect is unlikely to be zero.
- Economic significance says it is large enough to matter after costs, risk and capacity.
A strategy with t = 5 and an edge of 0.02% per trade may be worthless after costs, while a strategy with t = 1.5 but a strong economic rationale might deserve further study with small size. See Costs and Slippage in Backtests.
Beyond p values#
| Evidence | Why it helps |
|---|---|
| Economic rationale | A reason the edge should exist reduces the chance it is a fluke |
| Robustness across parameters | Results that survive small changes are more credible. See Robustness and Stress Testing |
| Consistency across markets and periods | Effects that appear widely are less likely to be chance |
| Out of sample and live performance | The strongest test |
| Low number of trials | Fewer tested variations means less data mining |
A practical checklist#
- How many independent trades or periods?
- What is the t statistic or confidence interval? See Confidence Intervals.
- How many variations were tried?
- Does it survive costs and realistic execution?
- Is there a sensible reason it should work?
- Does it hold out of sample?
Frequently asked questions#
What does statistically significant mean for a trading strategy?#
That its results would be unlikely if the strategy had no real edge, usually judged by a t statistic or p value.
How many years of data do I need to prove a strategy?#
It depends on the Sharpe ratio: roughly (2 / Sharpe)² years for a t statistic of 2, so 16 years for a Sharpe of 0.5 and 4 years for a Sharpe of 1.0.
Is statistical significance enough to trade a strategy?#
No. Results must also be economically meaningful after costs, robust, tested out of sample and adjusted for how many ideas were tried.
Next, learn to model relationships in Regression Analysis.
3 quick questions on this lesson. Get them all right to finish it.
Turn on JavaScript to take the quiz.
Mentioned in
- Law of Large NumbersMath and Statistics
- Confidence IntervalsMath and Statistics
- Statistical Power and Type I and II ErrorsMath and Statistics
- Outliers and Robust StatisticsMath and Statistics
- OptimizationMath and Statistics
- Quant Trading Learning PathStart Here