# Model Evaluation and Cross-Validation

> Standard cross validation leaks future data in time series. Learn time series splits, purging and embargo, combinatorial purged cross validation and good practice.

Source: https://learn.tradelabsai.com/machine-learning/cross-validation/  
Track: Machine Learning · Level: Advanced · Updated: 2026-10-03  
Publisher: TradeLabs AI (https://tradelabsai.com). Education, not financial advice.  
Cite as: TradeLabs Learn, "Model Evaluation and Cross-Validation", https://learn.tradelabsai.com/machine-learning/cross-validation/

Cross validation estimates how well a model will perform on data it has not seen, by repeatedly training on part of the data and testing on the rest. In most machine learning tutorials, the data is shuffled randomly into folds. For financial time series, that standard approach is dangerous: it trains on the future to predict the past, and overlapping labels leak information between folds. The result is validation scores that look far better than anything achievable live. Time aware validation methods fix this.

## Why random folds fail on market data

| Problem | Explanation |
|---|---|
| Training on the future | A random fold may train on 2024 and test on 2019 |
| Serial correlation | Neighbouring observations are similar, so test points have near copies in training |
| Overlapping labels | A 10 day forward return label on Monday overlaps the label on Tuesday |
| Regime information | Training on future data lets the model learn which regime came next |

## Time series split

The simplest correct method: always train on earlier data and test on later data. With an expanding window, each fold adds more history to training.

| Fold | Train | Test |
|---|---|---|
| 1 | 2015 to 2017 | 2018 |
| 2 | 2015 to 2018 | 2019 |
| 3 | 2015 to 2019 | 2020 |
| 4 | 2015 to 2020 | 2021 |

scikit-learn's `TimeSeriesSplit` implements this pattern. See [Walk-Forward Validation and Preventing Overfitting](https://learn.tradelabsai.com/machine-learning/walk-forward-validation/).

## Purging and embargo

When labels span future periods, training samples near the test period can contain information about it.

- **Purging:** remove training samples whose label periods overlap the test period.
- **Embargo:** also remove a short buffer of samples just after the test period, since serial correlation can leak information backwards.

These ideas were formalised by Marcos López de Prado in "Advances in Financial Machine Learning" (2018).

**Example: Purging overlapping labels**
Each sample's label is the return over the next 5 trading days. The test fold covers trading days 501 to 600. A training sample on day 497 has a label spanning days 498 to 502, which overlaps the first two test days. Keeping it would let the model learn part of the test outcome. Purging removes training samples from days 496 to 500, whose labels reach into the test period. With an embargo of 5 days, training samples from days 601 to 605 are also removed. The model loses 10 samples but gains an honest test.

## Combinatorial purged cross validation

Single train and test paths can be lucky or unlucky. Combinatorial purged cross validation (CPCV) splits data into groups and tests on many combinations of groups, with purging and embargo, producing a distribution of backtest results instead of one number. This helps assess how likely a strategy's performance is to be due to chance. See [Bootstrap and Permutation Tests](https://learn.tradelabsai.com/math/bootstrap-and-permutation-tests/).

## Nested validation for tuning

If you tune model settings using the same folds you report results on, the reported score is optimistic. Use an inner loop to choose settings and an outer loop, or a final untouched test period, to measure performance. See [Parameter Optimization](https://learn.tradelabsai.com/research/parameter-optimization/).

## Good practice

1. **Never shuffle** time series before splitting.
2. **Purge and embargo** when labels overlap.
3. **Keep a final holdout** period that is used once. See [In-Sample vs Out-of-Sample Testing](https://learn.tradelabsai.com/research/out-of-sample-testing/).
4. **Report the spread** of results across folds, not just the average.
5. **Count your trials:** each experiment increases the chance of a lucky result. See [P-Hacking and Multiple Testing](https://learn.tradelabsai.com/research/p-hacking-and-multiple-testing/).
6. **Test the trading rule with costs,** not only prediction metrics.

## Frequently asked questions

### Can I use k fold cross validation on stock data?

Standard shuffled k fold leaks future information. Use time series splits, or k fold with purging and embargo designed for financial data.

### What is purging in cross validation?

Removing training samples whose label periods overlap the test period, so the model cannot learn test outcomes indirectly.

### What is an embargo?

A buffer of samples removed after the test period to prevent leakage caused by serial correlation.

Next, learn the walk forward approach in detail in [Walk-Forward Validation and Preventing Overfitting](https://learn.tradelabsai.com/machine-learning/walk-forward-validation/).

## Continue learning

- Next lesson: [Walk-Forward Validation and Preventing Overfitting](https://learn.tradelabsai.com/machine-learning/walk-forward-validation/)
- Previous lesson: [Feature Engineering](https://learn.tradelabsai.com/machine-learning/feature-engineering/)
- Related: [Feature Engineering](https://learn.tradelabsai.com/machine-learning/feature-engineering/): Features are the inputs that give trading models a chance. Learn the main feature families, how to make them stationary and comparable, and how to avoid leakage.
- Related: [Walk-Forward Validation and Preventing Overfitting](https://learn.tradelabsai.com/machine-learning/walk-forward-validation/): Walk forward validation retrains a model on a rolling or expanding window and tests it on the next period, just as it would be used live. Learn setup and choices.
- Related: [In-Sample vs Out-of-Sample Testing](https://learn.tradelabsai.com/research/out-of-sample-testing/): Out of sample testing checks a strategy on data not used to build it. Learn train, validation and holdout splits, common mistakes and how to read results.
- Related: [Data Leakage](https://learn.tradelabsai.com/research/data-leakage/): Data leakage lets information from test data or the future slip into model training. Learn common leaks in trading and machine learning and how to prevent them.
- Related: [Overfitting and Curve Fitting](https://learn.tradelabsai.com/research/overfitting-and-curve-fitting/): Overfitting means a strategy fits noise instead of a real pattern. Learn the warning signs, why it happens, how to measure it and practical ways to avoid it.
- Related: [Supervised vs Unsupervised Learning](https://learn.tradelabsai.com/machine-learning/supervised-learning/): Supervised learning trains models on examples with known answers. Learn regression versus classification, how to define trading targets and labels, and key pitfalls.
