TradeLabs AILearn

Statistical Power and Type I and II Errors

Type I errors are false positives; type II errors are missed real effects. Learn how they apply to strategy testing, the trade off between them and power.

Intermediate3 min readUpdated 3 Oct 2026
Markdown
Lesson 15 of 46

Every statistical test can be wrong in two ways. A type I error, or false positive, means concluding that an effect exists when it does not, such as believing a strategy has an edge when its results were luck. A type II error, or false negative, means missing a real effect, such as discarding a genuinely profitable strategy because the test was not convincing. In trading research, both errors are costly, and the balance between them shapes how strict your testing should be.

The two errors#

Reality: no edgeReality: real edge
Test says edgeType I error (false positive), probability αCorrect (power = 1 minus β)
Test says no edgeCorrectType II error (false negative), probability β
  • α (alpha): the significance level, the accepted rate of false positives, often 5%.
  • β (beta): the probability of missing a real effect.
  • Power: 1 minus β, the probability of detecting a real effect when it exists.

Costs in trading#

ErrorTrading consequence
Type I (false positive)Trading a strategy with no edge: losses from costs and randomness, wasted capital and time
Type II (false negative)Discarding a real edge: missed profits

Most experienced researchers consider false positives more dangerous in trading, because markets are noisy, many ideas are tested and trading a fake edge loses real money. This is why quant firms set high bars for evidence. See P-Hacking and Multiple Testing.

The trade off#

Making a test stricter (lower α) reduces false positives but increases false negatives, for the same amount of data. The only way to reduce both is to get more or better data: more trades, longer histories or larger effects.

Statistical power#

Power depends on:

FactorEffect on power
Sample sizeMore data, more power
Effect sizeLarger edges are easier to detect
Noise (volatility)More noise, less power
Significance levelStricter α, less power

Base rates and false discoveries#

Even with α = 5%, if most ideas you test have no edge, many of your "significant" findings will be false.

Reducing errors in practice#

  1. Formulate hypotheses before testing, based on economic reasoning.
  2. Limit the number of variations tested, and record all tests.
  3. Use stricter thresholds when many ideas are tested, such as t statistics above 3.
  4. Use out of sample and walk forward tests. See In-Sample vs Out-of-Sample Testing and Walk-Forward Analysis.
  5. Paper trade or trade small before scaling up. See Moving From Paper to Live Trading.
  6. Collect more data to increase power.

Frequently asked questions#

What is a type I error?#

A false positive: concluding an effect exists, such as a trading edge, when it is actually due to chance.

What is a type II error?#

A false negative: failing to detect an effect that really exists, such as rejecting a strategy that truly has an edge.

Which error is worse in trading?#

Many researchers consider false positives worse, because trading a strategy with no real edge loses money, and testing many ideas makes false positives common.

Next, learn resampling methods in Bootstrap and Permutation Tests.

Check your understanding

3 quick questions on this lesson. Get them all right to finish it.

Turn on JavaScript to take the quiz.

Finished this lesson?Sign in to save your progress across devices.
Next lessonBootstrap and Permutation TestsBootstrap and permutation tests use resampling to measure uncertainty and test significance without strict assumptions. Learn how they work, examples and pitfalls.

Mentioned in