# Data Leakage

> Data leakage lets information from test data or the future slip into model training. Learn common leaks in trading and machine learning and how to prevent them.

Source: https://learn.tradelabsai.com/research/data-leakage/  
Track: Research and Backtesting · Level: Intermediate · Updated: 2026-10-03  
Publisher: TradeLabs AI (https://tradelabsai.com). Education, not financial advice.  
Cite as: TradeLabs Learn, "Data Leakage", https://learn.tradelabsai.com/research/data-leakage/

Data leakage happens when information that should not be available during model training or decision making slips in, making performance look far better than it will be in reality. Look ahead bias is one form of leakage, but leakage is broader and especially common in machine learning, where complex pipelines can mix training and test data in subtle ways. In trading, leakage turns ordinary models into apparent goldmines, and it is often discovered only after money is lost.

## Types of leakage

| Type | Example |
|---|---|
| Future information in features | Using a feature calculated with data from after the prediction time. See [Look-Ahead Bias](https://learn.tradelabsai.com/research/look-ahead-bias/) |
| Train test contamination | Randomly splitting time series so training includes data after test points |
| Overlapping labels | Labels covering periods that overlap between training and test sets |
| Preprocessing on all data | Scaling, filling missing values or selecting features using the full dataset |
| Target leakage | A feature that is effectively a version of the target |
| Repeated test set use | Tuning hyperparameters on the test set |
| Duplicate or near duplicate records | The same event in both training and test sets |

## Leakage in time series machine learning

Random k fold cross validation, common in general machine learning, leaks information in financial time series because observations close in time are correlated. Training on Tuesday and testing on Monday of the same week lets the model use information about the future.

**Example: Overlapping labels**
A model predicts 20 day forward returns using daily data. Each label spans the next 20 days. If the training set includes a sample from 1 March (label covering 1 to 28 March) and the test set includes 10 March, their labels overlap heavily. The model effectively sees part of the test outcome during training. Marcos López de Prado recommended "purging" training samples whose labels overlap the test period and adding an "embargo" gap after it. See [Model Evaluation and Cross-Validation](https://learn.tradelabsai.com/machine-learning/cross-validation/) and [Walk-Forward Validation and Preventing Overfitting](https://learn.tradelabsai.com/machine-learning/walk-forward-validation/).

## Preprocessing leaks

| Step | Leaky approach | Safe approach |
|---|---|---|
| Scaling | Fit scaler on the full dataset | Fit on training data only, apply to test |
| Missing values | Fill using the full sample mean | Use only past data |
| Feature selection | Pick features using all data | Select within training folds only |
| Outlier removal | Remove based on full sample statistics | Use rules or training statistics only |

## Target leakage examples in trading

- **Using the day's high or low** as a feature to predict whether the day closes up.
- **Using end of day volume** to predict intraday moves.
- **Using analyst upgrades dated to the day** when they were actually published after the market close.
- **Using an index's constituents** that were added because of the very performance being predicted.

## How to prevent leakage

1. **Map the timeline:** for every feature and label, know exactly when it becomes available.
2. **Split data in time order** with gaps (purging and embargo) where labels overlap.
3. **Build pipelines that fit transformations on training data only.**
4. **Keep a final holdout untouched.** See [In-Sample vs Out-of-Sample Testing](https://learn.tradelabsai.com/research/out-of-sample-testing/).
5. **Check suspicious results:** if a model is "too good", hunt for leaks before celebrating.
6. **Test with a deliberate delay:** shifting features later in time should hurt performance only modestly.
7. **Peer review code and data timing.**

## Leakage checklist for a new feature

- When is the raw data published?
- Is it revised later, and which version am I using? See [Point-in-Time and Survivorship-Free Data](https://learn.tradelabsai.com/programming/point-in-time-data/).
- Does any calculation use future rows?
- Could the feature encode the label?
- Are training and test samples independent in time?

## Frequently asked questions

### What is data leakage in trading models?

When information from the future or from the test set slips into model training or features, making performance look unrealistically good.

### How is data leakage different from look ahead bias?

Look ahead bias is a type of leakage involving future information; leakage also includes train test contamination, preprocessing on all data and target leakage.

### How do I prevent data leakage in machine learning for trading?

Split data in time order with purging and embargo gaps, fit preprocessing on training data only, track when each feature becomes available and keep a final holdout.

Next, learn how testing many ideas creates false discoveries in [P-Hacking and Multiple Testing](https://learn.tradelabsai.com/research/p-hacking-and-multiple-testing/).

## Continue learning

- Next lesson: [P-Hacking and Multiple Testing](https://learn.tradelabsai.com/research/p-hacking-and-multiple-testing/)
- Previous lesson: [Survivorship and Selection Bias](https://learn.tradelabsai.com/research/survivorship-and-selection-bias/)
- Related: [Survivorship and Selection Bias](https://learn.tradelabsai.com/research/survivorship-and-selection-bias/): Survivorship bias ignores failures; selection bias picks unrepresentative samples. Learn how both inflate backtests, fund returns and advice, and how to fix them.
- Related: [Look-Ahead Bias](https://learn.tradelabsai.com/research/look-ahead-bias/): Look ahead bias happens when a backtest uses information that was not available at the time. Learn common sources, real examples and how to prevent it in code.
- Related: [Model Evaluation and Cross-Validation](https://learn.tradelabsai.com/machine-learning/cross-validation/): Standard cross validation leaks future data in time series. Learn time series splits, purging and embargo, combinatorial purged cross validation and good practice.
- Related: [Walk-Forward Validation and Preventing Overfitting](https://learn.tradelabsai.com/machine-learning/walk-forward-validation/): Walk forward validation retrains a model on a rolling or expanding window and tests it on the next period, just as it would be used live. Learn setup and choices.
- Related: [Feature Engineering](https://learn.tradelabsai.com/machine-learning/feature-engineering/): Features are the inputs that give trading models a chance. Learn the main feature families, how to make them stationary and comparable, and how to avoid leakage.
- Related: [In-Sample vs Out-of-Sample Testing](https://learn.tradelabsai.com/research/out-of-sample-testing/): Out of sample testing checks a strategy on data not used to build it. Learn train, validation and holdout splits, common mistakes and how to read results.
