# Statistical Power and Type I and II Errors

> Type I errors are false positives; type II errors are missed real effects. Learn how they apply to strategy testing, the trade off between them and power.

Source: https://learn.tradelabsai.com/math/type-i-and-ii-errors/  
Track: Math and Statistics · Level: Intermediate · Updated: 2026-10-03  
Publisher: TradeLabs AI (https://tradelabsai.com). Education, not financial advice.  
Cite as: TradeLabs Learn, "Statistical Power and Type I and II Errors", https://learn.tradelabsai.com/math/type-i-and-ii-errors/

Every statistical test can be wrong in two ways. A type I error, or false positive, means concluding that an effect exists when it does not, such as believing a strategy has an edge when its results were luck. A type II error, or false negative, means missing a real effect, such as discarding a genuinely profitable strategy because the test was not convincing. In trading research, both errors are costly, and the balance between them shapes how strict your testing should be.

## The two errors

| | Reality: no edge | Reality: real edge |
|---|---|---|
| Test says edge | Type I error (false positive), probability α | Correct (power = 1 minus β) |
| Test says no edge | Correct | Type II error (false negative), probability β |

- **α (alpha):** the significance level, the accepted rate of false positives, often 5%.
- **β (beta):** the probability of missing a real effect.
- **Power:** 1 minus β, the probability of detecting a real effect when it exists.

## Costs in trading

| Error | Trading consequence |
|---|---|
| Type I (false positive) | Trading a strategy with no edge: losses from costs and randomness, wasted capital and time |
| Type II (false negative) | Discarding a real edge: missed profits |

Most experienced researchers consider false positives more dangerous in trading, because markets are noisy, many ideas are tested and trading a fake edge loses real money. This is why quant firms set high bars for evidence. See [P-Hacking and Multiple Testing](https://learn.tradelabsai.com/research/p-hacking-and-multiple-testing/).

## The trade off

Making a test stricter (lower α) reduces false positives but increases false negatives, for the same amount of data. The only way to reduce both is to get more or better data: more trades, longer histories or larger effects.

## Statistical power

Power depends on:

| Factor | Effect on power |
|---|---|
| Sample size | More data, more power |
| Effect size | Larger edges are easier to detect |
| Noise (volatility) | More noise, less power |
| Significance level | Stricter α, less power |

**Example: Power of a backtest**
A strategy truly earns 0.2% per trade with a standard deviation of 2%, an effect size of 0.1 standard deviations. To detect this with 80% power at α = 5% (one sided), you need roughly (1.645 + 0.84)² / 0.1² ≈ 620 trades. With only 100 trades, power is about 26%: three times out of four, a test would fail to confirm this real edge. Small, real edges are hard to prove. See [Sampling and Standard Error](https://learn.tradelabsai.com/math/sampling-and-standard-error/).

## Base rates and false discoveries

Even with α = 5%, if most ideas you test have no edge, many of your "significant" findings will be false.

**Example: How many discoveries are real?**
Suppose 10% of the strategies you test have a real edge, your tests have 50% power and α = 5%. Out of 1,000 ideas: 100 are real, and 50 of those pass; 900 are fake, and 45 of those pass by chance. Of the 95 that pass, 45 (about 47%) are false positives. A significant result is far from a guarantee. This is the logic behind concerns about false discoveries in finance research. See [Bayes' Theorem](https://learn.tradelabsai.com/math/bayes-theorem/).

## Reducing errors in practice

1. **Formulate hypotheses before testing,** based on economic reasoning.
2. **Limit the number of variations tested,** and record all tests.
3. **Use stricter thresholds** when many ideas are tested, such as t statistics above 3.
4. **Use out of sample and walk forward tests.** See [In-Sample vs Out-of-Sample Testing](https://learn.tradelabsai.com/research/out-of-sample-testing/) and [Walk-Forward Analysis](https://learn.tradelabsai.com/research/walk-forward-analysis/).
5. **Paper trade or trade small** before scaling up. See [Moving From Paper to Live Trading](https://learn.tradelabsai.com/start-here/paper-to-live-trading/).
6. **Collect more data** to increase power.

## Frequently asked questions

### What is a type I error?

A false positive: concluding an effect exists, such as a trading edge, when it is actually due to chance.

### What is a type II error?

A false negative: failing to detect an effect that really exists, such as rejecting a strategy that truly has an edge.

### Which error is worse in trading?

Many researchers consider false positives worse, because trading a strategy with no real edge loses money, and testing many ideas makes false positives common.

Next, learn resampling methods in [Bootstrap and Permutation Tests](https://learn.tradelabsai.com/math/bootstrap-and-permutation-tests/).

## Continue learning

- Next lesson: [Bootstrap and Permutation Tests](https://learn.tradelabsai.com/math/bootstrap-and-permutation-tests/)
- Previous lesson: [Hypothesis Testing and P-Values](https://learn.tradelabsai.com/math/hypothesis-testing-and-p-values/)
- Related: [Hypothesis Testing and P-Values](https://learn.tradelabsai.com/math/hypothesis-testing-and-p-values/): Hypothesis tests check whether results are likely due to chance. Learn null hypotheses, test statistics and p values, a strategy test and how p values mislead.
- Related: [P-Hacking and Multiple Testing](https://learn.tradelabsai.com/research/p-hacking-and-multiple-testing/): Testing many strategy variations guarantees some look good by chance. Learn how p hacking happens, how to adjust for multiple tests and the deflated Sharpe ratio.
- Related: [Statistical Significance in Trading](https://learn.tradelabsai.com/math/statistical-significance/): Statistical significance helps judge whether trading results reflect a real edge or luck. Learn the t statistic rule of thumb, sample size and multiple testing.
- Related: [Overfitting and Curve Fitting](https://learn.tradelabsai.com/research/overfitting-and-curve-fitting/): Overfitting means a strategy fits noise instead of a real pattern. Learn the warning signs, why it happens, how to measure it and practical ways to avoid it.
- Related: [Sampling and Standard Error](https://learn.tradelabsai.com/math/sampling-and-standard-error/): Standard error measures how much an estimate like a win rate or average return varies between samples. Learn the formulas and what they mean for backtests.
