TradeLabs AILearn

Bootstrap and Permutation Tests

Bootstrap and permutation tests use resampling to measure uncertainty and test significance without strict assumptions. Learn how they work, examples and pitfalls.

Intermediate3 min readUpdated 3 Oct 2026
Markdown
Lesson 16 of 46

Classical statistical tests often assume that returns are normal and independent, assumptions that market data frequently violates. Resampling methods offer an alternative. The bootstrap estimates uncertainty by repeatedly resampling your own data. Permutation tests check significance by shuffling data to see how often random arrangements produce results as good as yours. Both are flexible, intuitive and widely used to stress test trading strategies.

The bootstrap#

The bootstrap, introduced by Bradley Efron in 1979, treats your sample as a stand in for the population:

  1. Draw a new sample of the same size from your data, with replacement.
  2. Calculate the statistic (mean return, Sharpe ratio, maximum drawdown).
  3. Repeat thousands of times.
  4. Use the distribution of the statistic to estimate standard errors and confidence intervals.

The block bootstrap#

Simple bootstrapping assumes observations are independent. Market returns show volatility clustering and sometimes autocorrelation, so a block bootstrap resamples blocks of consecutive returns (for example, 20 days at a time) to preserve short term dependence. Variants include the stationary bootstrap of Politis and Romano. See Autocorrelation and Partial Autocorrelation and GARCH.

Bootstrapping drawdowns#

Maximum drawdown depends on the order of returns, so a single backtest gives just one path. Resampling returns or trades produces many alternative paths, showing how bad drawdowns could plausibly be. This helps set realistic expectations and position sizes. See Maximum Drawdown and Monte Carlo Simulation.

Permutation tests#

A permutation test asks: if the signal had no real relationship with future returns, how often would random shuffling produce results as good as mine?

  1. Calculate the strategy's actual performance.
  2. Shuffle the link between signals and returns (for example, randomly reorder the signals in time).
  3. Recalculate performance for each shuffle.
  4. Repeat thousands of times.
  5. p value = share of shuffles that perform as well as or better than the real strategy.

Advantages#

  • Few distribution assumptions.
  • Work for complex statistics such as Sharpe ratios and drawdowns.
  • Intuitive and easy to explain.
  • Easy to implement in Python. See Python for Trading.

Pitfalls#

PitfallExplanation
Ignoring dependenceSimple bootstraps understate uncertainty with clustered data
Small samplesResampling cannot add information that is not in the data
Missing rare eventsIf a crash is not in your sample, the bootstrap will not create it. See Fat Tails
Data miningResampling does not fix bias from testing many strategies. See P-Hacking and Multiple Testing
Shuffling destroys structurePermutations must preserve the right features (such as trend) to be a fair comparison

White's Reality Check (2000) and Hansen's Superior Predictive Ability test use bootstrap methods to test whether the best of many strategies is truly better than a benchmark, addressing data snooping directly.

Frequently asked questions#

What is bootstrapping in trading?#

A resampling method that repeatedly draws samples from historical returns or trades to estimate the uncertainty of statistics like Sharpe ratios and drawdowns.

What is a permutation test?#

A test that shuffles the link between signals and outcomes many times to see how often random arrangements match the real result, giving a p value.

When should I use bootstrap methods?#

When returns are non normal, statistics are complex, or you want to see the range of possible outcomes such as drawdowns.

Next, learn how to handle extreme values in Outliers and Robust Statistics.

Check your understanding

3 quick questions on this lesson. Get them all right to finish it.

Turn on JavaScript to take the quiz.

Finished this lesson?Sign in to save your progress across devices.
Next lessonOutliers and Robust StatisticsOutliers can distort averages, correlations and backtests. Learn how to detect them, robust measures like the median and MAD, winsorising and when they matter.

Mentioned in