Backtest Reproducibility
A reproducible backtest gives the same results every time from the same code and data. Learn version control, data snapshots, research logs and good habits.
A backtest you cannot reproduce is a backtest you cannot trust. If running the same strategy next month produces different numbers, or nobody remembers which data and parameters produced last year's impressive result, you cannot know whether the edge was real, a bug or a lucky choice. Reproducibility means that the same code, data and settings always produce the same results, and that every step from raw data to final report is documented. It is standard practice at professional quant firms and easy for individual traders to adopt.
Why results fail to reproduce#
| Cause | Example |
|---|---|
| Changing data | Vendor revisions, adjusted prices recalculated, new corporate actions |
| Code changes | A bug fix or tweak not tracked |
| Undocumented parameters | Settings changed in a notebook and forgotten |
| Randomness | Monte Carlo or machine learning without fixed random seeds |
| Library versions | Package updates changing behaviour |
| Manual steps | Hand edited spreadsheets or one off data fixes |
| Cherry picking | Only the best run was saved. See P-Hacking and Multiple Testing |
The building blocks#
| Practice | How |
|---|---|
| Version control | Keep all code in Git, with meaningful commit messages |
| Data snapshots | Store raw data and record exactly which version each backtest used. See Data Versioning, Lineage and Schemas |
| Configuration files | Keep parameters in config files, not scattered in code |
| Fixed random seeds | Make simulations repeatable |
| Environment pinning | Record library versions (for example, a requirements file or lock file) |
| Automated pipelines | Scripts that run from raw data to results without manual steps |
| Research log | Record every experiment, its purpose and outcome |
A research log entry#
Notebooks vs scripts#
Jupyter notebooks are excellent for exploration but easy to run out of order, leaving hidden state. Once an idea looks promising, move the logic into tested, version controlled scripts or modules, and rerun everything from scratch to confirm results. See Python for Trading.
Testing your code#
- Unit tests for indicator calculations and signal logic.
- Known answer tests: check calculations against hand computed examples.
- Regression tests: confirm that results do not change unexpectedly after code updates.
- Look ahead checks: shift data forward and confirm results change. See Look-Ahead Bias.
Reproducibility in live trading#
The same principles apply after deployment: log every signal, order and fill, and keep the code version used for each trading day. This makes it possible to explain any trade and to compare live behaviour with backtests. See Logging, Audit Trails and Incident Response and Monitoring Positions, P&L and Risk.
Benefits#
- Catch bugs before they cost money.
- Honest evaluation: know how many ideas were tried.
- Collaboration: others can verify and build on your work.
- Faster iteration: rerun everything with one command.
- Regulatory and investor confidence for professional managers.
Frequently asked questions#
What is backtest reproducibility?#
The ability to get exactly the same backtest results from the same code, data and settings, with every step documented.
How do I make my backtests reproducible?#
Use version control, store data snapshots, keep parameters in configuration files, fix random seeds, pin library versions and keep a research log.
Why do backtest results change over time?#
Common causes include revised or re adjusted data, untracked code changes, library updates and randomness without fixed seeds.
Next, learn what traders mean by alpha in What Is Alpha?.
3 quick questions on this lesson. Get them all right to finish it.
Turn on JavaScript to take the quiz.
Mentioned in
- Backtesting MethodologyResearch and Backtesting
- Robustness and Stress TestingResearch and Backtesting
- Event-Driven vs Vectorized BacktestingResearch and Backtesting
- Corporate Actions, Delistings and Rolls in BacktestsResearch and Backtesting
- Outliers and Robust StatisticsMath and Statistics
- Developing, Testing and Monitoring AlgorithmsAlgorithmic Trading