# Statistical Significance in Trading

> Statistical significance helps judge whether trading results reflect a real edge or luck. Learn the t statistic rule of thumb, sample size and multiple testing.

Source: https://learn.tradelabsai.com/math/statistical-significance/  
Track: Math and Statistics · Level: Intermediate · Updated: 2026-10-03  
Publisher: TradeLabs AI (https://tradelabsai.com). Education, not financial advice.  
Cite as: TradeLabs Learn, "Statistical Significance in Trading", https://learn.tradelabsai.com/math/statistical-significance/

A strategy that made money in a backtest or a few months of live trading may have a real edge, or it may have been lucky. Statistical significance is a way of judging how unlikely the results would be if there were no edge at all. It is not proof, and it can be badly misused, but it provides a disciplined first filter. This lesson brings together the ideas from earlier lessons into practical rules for judging trading results.

## The t statistic as a quick guide

For average returns:

```
t = mean return / (standard deviation / √n)
```

For an annualised Sharpe ratio measured over T years, a convenient approximation is:

```
t ≈ Sharpe ratio × √T
```

| t statistic | Rough interpretation (single test, no data mining) |
|---|---|
| Below 1 | Indistinguishable from noise |
| About 2 | Conventionally "significant" at 5% |
| Above 3 | Strong evidence; recommended for new factors after many have been tested |

**Example: How long to prove a Sharpe ratio?**
To reach t = 2 with a true Sharpe ratio of 0.5, you need about (2 / 0.5)² = 16 years of data. With a Sharpe of 1.0, about 4 years. With a Sharpe of 2.0, about 1 year. Good but not spectacular strategies take a long time to prove statistically, which is why traders combine statistics with economic reasoning. See [Sharpe Ratio](https://learn.tradelabsai.com/portfolio/sharpe-ratio/).

## Sample size matters more than you think

Statistical significance depends heavily on the number of independent observations. Strategies that trade rarely, such as a few times a year, may never produce enough trades to be statistically convincing within a reasonable time. High frequency strategies accumulate observations quickly, which partly explains why their edges can be measured more precisely. See [Sampling and Standard Error](https://learn.tradelabsai.com/math/sampling-and-standard-error/).

## Multiple testing changes everything

If you test many ideas, some will look significant by chance. With 20 independent tests of strategies with no edge, the chance that at least one shows p < 0.05 is 1 minus 0.95^20 ≈ 64%.

Ways to adjust:

| Method | Approach |
|---|---|
| Bonferroni correction | Divide α by the number of tests; strict |
| Holm and Benjamini Hochberg | Less strict; control the false discovery rate |
| Higher t thresholds | Such as t > 3, as suggested by Harvey, Liu and Zhu (2016) |
| Deflated Sharpe ratio | Adjusts the Sharpe ratio for the number of trials and non normal returns (Bailey and López de Prado) |
| Out of sample testing | Test on data not used to develop the strategy. See [In-Sample vs Out-of-Sample Testing](https://learn.tradelabsai.com/research/out-of-sample-testing/) |

See [P-Hacking and Multiple Testing](https://learn.tradelabsai.com/research/p-hacking-and-multiple-testing/).

## Significance is not importance

- **Statistical significance** says an effect is unlikely to be zero.
- **Economic significance** says it is large enough to matter after costs, risk and capacity.

A strategy with t = 5 and an edge of 0.02% per trade may be worthless after costs, while a strategy with t = 1.5 but a strong economic rationale might deserve further study with small size. See [Costs and Slippage in Backtests](https://learn.tradelabsai.com/research/costs-and-slippage-in-backtests/).

## Beyond p values

| Evidence | Why it helps |
|---|---|
| Economic rationale | A reason the edge should exist reduces the chance it is a fluke |
| Robustness across parameters | Results that survive small changes are more credible. See [Robustness and Stress Testing](https://learn.tradelabsai.com/research/robustness-and-stress-testing/) |
| Consistency across markets and periods | Effects that appear widely are less likely to be chance |
| Out of sample and live performance | The strongest test |
| Low number of trials | Fewer tested variations means less data mining |

## A practical checklist

1. **How many independent trades or periods?**
2. **What is the t statistic or confidence interval?** See [Confidence Intervals](https://learn.tradelabsai.com/math/confidence-intervals/).
3. **How many variations were tried?**
4. **Does it survive costs and realistic execution?**
5. **Is there a sensible reason it should work?**
6. **Does it hold out of sample?**

## Frequently asked questions

### What does statistically significant mean for a trading strategy?

That its results would be unlikely if the strategy had no real edge, usually judged by a t statistic or p value.

### How many years of data do I need to prove a strategy?

It depends on the Sharpe ratio: roughly (2 / Sharpe)² years for a t statistic of 2, so 16 years for a Sharpe of 0.5 and 4 years for a Sharpe of 1.0.

### Is statistical significance enough to trade a strategy?

No. Results must also be economically meaningful after costs, robust, tested out of sample and adjusted for how many ideas were tried.

Next, learn to model relationships in [Regression Analysis](https://learn.tradelabsai.com/math/regression-analysis/).

## Continue learning

- Next lesson: [Regression Analysis](https://learn.tradelabsai.com/math/regression-analysis/)
- Previous lesson: [Outliers and Robust Statistics](https://learn.tradelabsai.com/math/outliers-and-robust-statistics/)
- Related: [Outliers and Robust Statistics](https://learn.tradelabsai.com/math/outliers-and-robust-statistics/): Outliers can distort averages, correlations and backtests. Learn how to detect them, robust measures like the median and MAD, winsorising and when they matter.
- Related: [Hypothesis Testing and P-Values](https://learn.tradelabsai.com/math/hypothesis-testing-and-p-values/): Hypothesis tests check whether results are likely due to chance. Learn null hypotheses, test statistics and p values, a strategy test and how p values mislead.
- Related: [P-Hacking and Multiple Testing](https://learn.tradelabsai.com/research/p-hacking-and-multiple-testing/): Testing many strategy variations guarantees some look good by chance. Learn how p hacking happens, how to adjust for multiple tests and the deflated Sharpe ratio.
- Related: [Sharpe Ratio](https://learn.tradelabsai.com/portfolio/sharpe-ratio/): The Sharpe ratio measures return per unit of risk. Learn the formula, how to annualise it, what counts as a good Sharpe ratio, its limitations and common mistakes.
- Related: [Sampling and Standard Error](https://learn.tradelabsai.com/math/sampling-and-standard-error/): Standard error measures how much an estimate like a win rate or average return varies between samples. Learn the formulas and what they mean for backtests.
- Related: [In-Sample vs Out-of-Sample Testing](https://learn.tradelabsai.com/research/out-of-sample-testing/): Out of sample testing checks a strategy on data not used to build it. Learn train, validation and holdout splits, common mistakes and how to read results.
