In-Sample vs Out-of-Sample Testing
Out of sample testing checks a strategy on data not used to build it. Learn train, validation and holdout splits, common mistakes and how to read results.
A strategy tuned on historical data will always look good on that same data. The real question is how it performs on data it has never seen. Out of sample testing answers this by setting aside part of the data during development and using it only to evaluate the final strategy. It is one of the most important defences against overfitting and one of the most commonly misused.
In sample vs out of sample#
| In sample | Out of sample | |
|---|---|---|
| Used for | Developing rules, choosing parameters | Evaluating the finished strategy |
| Performance | Usually optimistic | A more honest estimate |
| How often to use | Freely during development | Ideally once |
A common split#
| Segment | Share of data | Purpose |
|---|---|---|
| Training (development) | About 60% | Build and tune the strategy |
| Validation | About 20% | Compare candidate versions and catch overfitting |
| Holdout (test) | About 20% | Final, one time evaluation |
For time series, segments should be in time order: train on earlier data, validate and test on later data. Random splits leak future information into the past. See Data Leakage.
The holdout must stay untouched#
The most common mistake is peeking: testing on the holdout, seeing poor results, adjusting the strategy and testing again. Each time, the holdout becomes part of the development data, and its results become optimistic. If you use the holdout more than once, treat its results as in sample. See P-Hacking and Multiple Testing.
Other forms of out of sample testing#
| Method | How it works | Lesson |
|---|---|---|
| Walk forward analysis | Repeatedly train on past windows and test on the next period | Walk-Forward Analysis |
| Cross market tests | Apply the strategy unchanged to other markets | |
| Cross period tests | Test in earlier historical periods not used | |
| Time series cross validation | Several train and test splits in time order | Model Evaluation and Cross-Validation |
| Paper trading and live incubation | True out of sample, in real time | Paper Trading |
Limits#
- Short out of sample periods produce noisy estimates. See Sampling and Standard Error.
- Regime differences: the out of sample period may be unusual. See Structural Breaks and Regime Changes.
- Indirect leakage: knowledge of what happened in recent years can influence strategy design even if data is formally held out.
- Multiple strategies: if you test many strategies on the same holdout and pick the best, the holdout is no longer clean.
Best practices#
- Decide the split before any analysis.
- Keep the holdout locked until the strategy is final.
- Use walk forward or several out of sample tests when data allows.
- Report both in and out of sample results.
- Expect degradation and judge whether the out of sample result is still worth trading.
- Follow up with live incubation. See Moving From Paper to Live Trading.
Frequently asked questions#
What is out of sample testing?#
Testing a strategy on data that was not used to develop or tune it, to get an honest estimate of future performance.
How much data should be out of sample?#
Commonly 20% to 30% of the history, in time order, while ensuring both segments are long enough and cover varied market conditions.
Why did my strategy fail out of sample?#
Usually because of overfitting during development, data leakage or a change in market conditions between the two periods.
Next, learn a rolling version of this idea in Walk-Forward Analysis.
3 quick questions on this lesson. Get them all right to finish it.
Turn on JavaScript to take the quiz.
Mentioned in
- The Trading Research ProcessResearch and Backtesting
- Backtesting MethodologyResearch and Backtesting
- Historical Data for BacktestingResearch and Backtesting
- Robustness and Stress TestingResearch and Backtesting
- Signal DiscoveryResearch and Backtesting
- Quant Trading Learning PathStart Here