Data Versioning, Lineage and Schemas
Data versioning tracks exactly which data, code and settings produced each backtest. Learn snapshots, hashes, tools like git and DVC, and a simple workflow.
Six months after a promising backtest, can you reproduce it exactly? Data vendors revise history, pipelines change, adjustment factors update with each dividend and code evolves. Without versioning, a strategy's results can drift for reasons nobody can trace, and it becomes impossible to tell whether a change in performance comes from the market or from the data. Data versioning means recording exactly which version of the data, code and parameters produced each result, so any result can be rebuilt and compared.
What needs versioning#
| Item | Why it changes | Tool |
|---|---|---|
| Code | Bug fixes and new features | git |
| Parameters and configuration | Tuning and experiments | git, config files |
| Raw data | Vendor corrections, new downloads | Snapshots, DVC, object storage versions |
| Processed data | Pipeline changes | Rebuild from versioned raw data and code |
| Environment | Library upgrades change results | Lock files, containers. See Docker and Kubernetes |
| Results | Each run's output | Experiment logs |
Approaches#
| Approach | How it works | Good for |
|---|---|---|
| Dated snapshots | Save each download in a folder named by date | Simple, small to medium datasets |
| Content hashes | Fingerprint each file; any change produces a new hash | Detecting silent changes |
| Append only tables | Never overwrite; add rows with a version or timestamp | Databases. See Point-in-Time and Survivorship-Free Data |
| Data version control tools | DVC, lakeFS or similar track large files alongside git | Larger research teams |
| Table formats with time travel | Formats such as Delta Lake or Apache Iceberg keep historical versions | Large data lakes |
A simple workflow for individuals#
- Keep raw downloads in dated folders and never edit them.
- Compute a hash (such as SHA 256) for each dataset used in a backtest.
- Commit code and configuration to git before each important run.
- Log each run with the git commit ID, data hashes, parameters and results.
- Pin library versions in a requirements or lock file.
import hashlib, json, subprocess
def file_hash(path):
h = hashlib.sha256()
with open(path, "rb") as f:
for chunk in iter(lambda: f.read(1 << 20), b""):
h.update(chunk)
return h.hexdigest()[:16]
run_record = {
"commit": subprocess.check_output(["git", "rev-parse", "HEAD"]).decode().strip(),
"data": {"prices.parquet": file_hash("prices.parquet")},
"params": {"fast": 20, "slow": 50},
}
print(json.dumps(run_record))
Versioning and point in time data#
Versioning answers "what did my dataset look like when I ran this?" Point in time data answers "what was known in the market on that date?" Both are needed: point in time data protects against look ahead bias, and versioning protects reproducibility. See Backtest Reproducibility.
Storage considerations#
Full copies of large datasets can become expensive. Options include storing only changed partitions, keeping compressed columnar files and pruning old snapshots that no result depends on. Cloud object storage with versioning enabled is a simple, low maintenance choice. See Data Storage, Compression and Caching.
Frequently asked questions#
What is data versioning?#
Recording which exact version of data was used for each analysis, so results can be reproduced and changes in data can be detected.
Why do backtest results change when I rerun them?#
Data revisions, changed adjustment factors, code edits and library upgrades can all change results; versioning each one shows which caused the change.
What tools are used for data versioning?#
git for code, dated snapshots and hashes for simple setups, and tools such as DVC, lakeFS, Delta Lake or Apache Iceberg for larger datasets.
Next, learn where and how to store market data efficiently in Data Storage, Compression and Caching.
3 quick questions on this lesson. Get them all right to finish it.
Turn on JavaScript to take the quiz.
Mentioned in
- Python for TradingProgramming and Data
- Cleaning Market DataProgramming and Data
- Market Data ReplayProgramming and Data
- Historical Data for BacktestingResearch and Backtesting
- Developing, Testing and Monitoring AlgorithmsAlgorithmic Trading
- Docker and KubernetesTrading Infrastructure