Model Evaluation and Cross-Validation
Standard cross validation leaks future data in time series. Learn time series splits, purging and embargo, combinatorial purged cross validation and good practice.
Cross validation estimates how well a model will perform on data it has not seen, by repeatedly training on part of the data and testing on the rest. In most machine learning tutorials, the data is shuffled randomly into folds. For financial time series, that standard approach is dangerous: it trains on the future to predict the past, and overlapping labels leak information between folds. The result is validation scores that look far better than anything achievable live. Time aware validation methods fix this.
Why random folds fail on market data#
| Problem | Explanation |
|---|---|
| Training on the future | A random fold may train on 2024 and test on 2019 |
| Serial correlation | Neighbouring observations are similar, so test points have near copies in training |
| Overlapping labels | A 10 day forward return label on Monday overlaps the label on Tuesday |
| Regime information | Training on future data lets the model learn which regime came next |
Time series split#
The simplest correct method: always train on earlier data and test on later data. With an expanding window, each fold adds more history to training.
| Fold | Train | Test |
|---|---|---|
| 1 | 2015 to 2017 | 2018 |
| 2 | 2015 to 2018 | 2019 |
| 3 | 2015 to 2019 | 2020 |
| 4 | 2015 to 2020 | 2021 |
scikit-learn's TimeSeriesSplit implements this pattern. See Walk-Forward Validation and Preventing Overfitting.
Purging and embargo#
When labels span future periods, training samples near the test period can contain information about it.
- Purging: remove training samples whose label periods overlap the test period.
- Embargo: also remove a short buffer of samples just after the test period, since serial correlation can leak information backwards.
These ideas were formalised by Marcos López de Prado in "Advances in Financial Machine Learning" (2018).
Combinatorial purged cross validation#
Single train and test paths can be lucky or unlucky. Combinatorial purged cross validation (CPCV) splits data into groups and tests on many combinations of groups, with purging and embargo, producing a distribution of backtest results instead of one number. This helps assess how likely a strategy's performance is to be due to chance. See Bootstrap and Permutation Tests.
Nested validation for tuning#
If you tune model settings using the same folds you report results on, the reported score is optimistic. Use an inner loop to choose settings and an outer loop, or a final untouched test period, to measure performance. See Parameter Optimization.
Good practice#
- Never shuffle time series before splitting.
- Purge and embargo when labels overlap.
- Keep a final holdout period that is used once. See In-Sample vs Out-of-Sample Testing.
- Report the spread of results across folds, not just the average.
- Count your trials: each experiment increases the chance of a lucky result. See P-Hacking and Multiple Testing.
- Test the trading rule with costs, not only prediction metrics.
Frequently asked questions#
Can I use k fold cross validation on stock data?#
Standard shuffled k fold leaks future information. Use time series splits, or k fold with purging and embargo designed for financial data.
What is purging in cross validation?#
Removing training samples whose label periods overlap the test period, so the model cannot learn test outcomes indirectly.
What is an embargo?#
A buffer of samples removed after the test period to prevent leakage caused by serial correlation.
Next, learn the walk forward approach in detail in Walk-Forward Validation and Preventing Overfitting.
3 quick questions on this lesson. Get them all right to finish it.
Turn on JavaScript to take the quiz.
Mentioned in
- Machine Learning in TradingMachine Learning
- Regression and Classification ModelsMachine Learning
- Random Forests and Gradient BoostingMachine Learning
- Feature EngineeringMachine Learning
- Regression AnalysisMath and Statistics