Bootstrap and Permutation Tests
Bootstrap and permutation tests use resampling to measure uncertainty and test significance without strict assumptions. Learn how they work, examples and pitfalls.
Classical statistical tests often assume that returns are normal and independent, assumptions that market data frequently violates. Resampling methods offer an alternative. The bootstrap estimates uncertainty by repeatedly resampling your own data. Permutation tests check significance by shuffling data to see how often random arrangements produce results as good as yours. Both are flexible, intuitive and widely used to stress test trading strategies.
The bootstrap#
The bootstrap, introduced by Bradley Efron in 1979, treats your sample as a stand in for the population:
- Draw a new sample of the same size from your data, with replacement.
- Calculate the statistic (mean return, Sharpe ratio, maximum drawdown).
- Repeat thousands of times.
- Use the distribution of the statistic to estimate standard errors and confidence intervals.
The block bootstrap#
Simple bootstrapping assumes observations are independent. Market returns show volatility clustering and sometimes autocorrelation, so a block bootstrap resamples blocks of consecutive returns (for example, 20 days at a time) to preserve short term dependence. Variants include the stationary bootstrap of Politis and Romano. See Autocorrelation and Partial Autocorrelation and GARCH.
Bootstrapping drawdowns#
Maximum drawdown depends on the order of returns, so a single backtest gives just one path. Resampling returns or trades produces many alternative paths, showing how bad drawdowns could plausibly be. This helps set realistic expectations and position sizes. See Maximum Drawdown and Monte Carlo Simulation.
Permutation tests#
A permutation test asks: if the signal had no real relationship with future returns, how often would random shuffling produce results as good as mine?
- Calculate the strategy's actual performance.
- Shuffle the link between signals and returns (for example, randomly reorder the signals in time).
- Recalculate performance for each shuffle.
- Repeat thousands of times.
- p value = share of shuffles that perform as well as or better than the real strategy.
Advantages#
- Few distribution assumptions.
- Work for complex statistics such as Sharpe ratios and drawdowns.
- Intuitive and easy to explain.
- Easy to implement in Python. See Python for Trading.
Pitfalls#
| Pitfall | Explanation |
|---|---|
| Ignoring dependence | Simple bootstraps understate uncertainty with clustered data |
| Small samples | Resampling cannot add information that is not in the data |
| Missing rare events | If a crash is not in your sample, the bootstrap will not create it. See Fat Tails |
| Data mining | Resampling does not fix bias from testing many strategies. See P-Hacking and Multiple Testing |
| Shuffling destroys structure | Permutations must preserve the right features (such as trend) to be a fair comparison |
Related tools#
White's Reality Check (2000) and Hansen's Superior Predictive Ability test use bootstrap methods to test whether the best of many strategies is truly better than a benchmark, addressing data snooping directly.
Frequently asked questions#
What is bootstrapping in trading?#
A resampling method that repeatedly draws samples from historical returns or trades to estimate the uncertainty of statistics like Sharpe ratios and drawdowns.
What is a permutation test?#
A test that shuffles the link between signals and outcomes many times to see how often random arrangements match the real result, giving a p value.
When should I use bootstrap methods?#
When returns are non normal, statistics are complex, or you want to see the range of possible outcomes such as drawdowns.
Next, learn how to handle extreme values in Outliers and Robust Statistics.
3 quick questions on this lesson. Get them all right to finish it.
Turn on JavaScript to take the quiz.
Mentioned in
- Sampling and Standard ErrorMath and Statistics
- Central Limit TheoremMath and Statistics
- Statistical Power and Type I and II ErrorsMath and Statistics
- Probability Distributions ExplainedMath and Statistics
- White Noise and Random WalksMath and Statistics
- P-Hacking and Multiple TestingResearch and Backtesting