# In-Sample vs Out-of-Sample Testing

> Out of sample testing checks a strategy on data not used to build it. Learn train, validation and holdout splits, common mistakes and how to read results.

Source: https://learn.tradelabsai.com/research/out-of-sample-testing/  
Track: Research and Backtesting · Level: Intermediate · Updated: 2026-10-03  
Publisher: TradeLabs AI (https://tradelabsai.com). Education, not financial advice.  
Cite as: TradeLabs Learn, "In-Sample vs Out-of-Sample Testing", https://learn.tradelabsai.com/research/out-of-sample-testing/

A strategy tuned on historical data will always look good on that same data. The real question is how it performs on data it has never seen. Out of sample testing answers this by setting aside part of the data during development and using it only to evaluate the final strategy. It is one of the most important defences against overfitting and one of the most commonly misused.

## In sample vs out of sample

| | In sample | Out of sample |
|---|---|---|
| Used for | Developing rules, choosing parameters | Evaluating the finished strategy |
| Performance | Usually optimistic | A more honest estimate |
| How often to use | Freely during development | Ideally once |

## A common split

| Segment | Share of data | Purpose |
|---|---|---|
| Training (development) | About 60% | Build and tune the strategy |
| Validation | About 20% | Compare candidate versions and catch overfitting |
| Holdout (test) | About 20% | Final, one time evaluation |

For time series, segments should be in time order: train on earlier data, validate and test on later data. Random splits leak future information into the past. See [Data Leakage](https://learn.tradelabsai.com/research/data-leakage/).

**Example: Reading out of sample results**
A strategy shows a Sharpe ratio of 1.8 in sample (2008 to 2018) and 0.7 out of sample (2019 to 2023). A drop is normal; the question is whether 0.7 is still attractive after costs and whether the out of sample period is long enough to trust. If out of sample performance had been minus 0.2, the strategy would likely be overfit or its edge gone. A Sharpe ratio close to the in sample value would be encouraging, but rare. See [Statistical Significance in Trading](https://learn.tradelabsai.com/math/statistical-significance/).

## The holdout must stay untouched

The most common mistake is peeking: testing on the holdout, seeing poor results, adjusting the strategy and testing again. Each time, the holdout becomes part of the development data, and its results become optimistic. If you use the holdout more than once, treat its results as in sample. See [P-Hacking and Multiple Testing](https://learn.tradelabsai.com/research/p-hacking-and-multiple-testing/).

## Other forms of out of sample testing

| Method | How it works | Lesson |
|---|---|---|
| Walk forward analysis | Repeatedly train on past windows and test on the next period | [Walk-Forward Analysis](https://learn.tradelabsai.com/research/walk-forward-analysis/) |
| Cross market tests | Apply the strategy unchanged to other markets | |
| Cross period tests | Test in earlier historical periods not used | |
| Time series cross validation | Several train and test splits in time order | [Model Evaluation and Cross-Validation](https://learn.tradelabsai.com/machine-learning/cross-validation/) |
| Paper trading and live incubation | True out of sample, in real time | [Paper Trading](https://learn.tradelabsai.com/start-here/paper-trading/) |

## Limits

- **Short out of sample periods** produce noisy estimates. See [Sampling and Standard Error](https://learn.tradelabsai.com/math/sampling-and-standard-error/).
- **Regime differences:** the out of sample period may be unusual. See [Structural Breaks and Regime Changes](https://learn.tradelabsai.com/math/regime-changes/).
- **Indirect leakage:** knowledge of what happened in recent years can influence strategy design even if data is formally held out.
- **Multiple strategies:** if you test many strategies on the same holdout and pick the best, the holdout is no longer clean.

## Best practices

1. **Decide the split before any analysis.**
2. **Keep the holdout locked** until the strategy is final.
3. **Use walk forward or several out of sample tests** when data allows.
4. **Report both in and out of sample results.**
5. **Expect degradation** and judge whether the out of sample result is still worth trading.
6. **Follow up with live incubation.** See [Moving From Paper to Live Trading](https://learn.tradelabsai.com/start-here/paper-to-live-trading/).

## Frequently asked questions

### What is out of sample testing?

Testing a strategy on data that was not used to develop or tune it, to get an honest estimate of future performance.

### How much data should be out of sample?

Commonly 20% to 30% of the history, in time order, while ensuring both segments are long enough and cover varied market conditions.

### Why did my strategy fail out of sample?

Usually because of overfitting during development, data leakage or a change in market conditions between the two periods.

Next, learn a rolling version of this idea in [Walk-Forward Analysis](https://learn.tradelabsai.com/research/walk-forward-analysis/).

## Continue learning

- Next lesson: [Walk-Forward Analysis](https://learn.tradelabsai.com/research/walk-forward-analysis/)
- Previous lesson: [Historical Data for Backtesting](https://learn.tradelabsai.com/research/historical-data-for-backtesting/)
- Related: [Historical Data for Backtesting](https://learn.tradelabsai.com/research/historical-data-for-backtesting/): Backtests are only as good as their data. Learn data types and sources, quality checks, corporate action adjustments and how to avoid survivorship traps.
- Related: [Walk-Forward Analysis](https://learn.tradelabsai.com/research/walk-forward-analysis/): Walk forward analysis repeatedly optimises a strategy on past data and tests it on the next period. Learn how it works, window choices, efficiency ratios and limits.
- Related: [Overfitting and Curve Fitting](https://learn.tradelabsai.com/research/overfitting-and-curve-fitting/): Overfitting means a strategy fits noise instead of a real pattern. Learn the warning signs, why it happens, how to measure it and practical ways to avoid it.
- Related: [Model Evaluation and Cross-Validation](https://learn.tradelabsai.com/machine-learning/cross-validation/): Standard cross validation leaks future data in time series. Learn time series splits, purging and embargo, combinatorial purged cross validation and good practice.
- Related: [Data Leakage](https://learn.tradelabsai.com/research/data-leakage/): Data leakage lets information from test data or the future slip into model training. Learn common leaks in trading and machine learning and how to prevent them.
- Related: [P-Hacking and Multiple Testing](https://learn.tradelabsai.com/research/p-hacking-and-multiple-testing/): Testing many strategy variations guarantees some look good by chance. Learn how p hacking happens, how to adjust for multiple tests and the deflated Sharpe ratio.
