# Data Versioning, Lineage and Schemas

> Data versioning tracks exactly which data, code and settings produced each backtest. Learn snapshots, hashes, tools like git and DVC, and a simple workflow.

Source: https://learn.tradelabsai.com/programming/data-versioning/  
Track: Programming and Data · Level: Advanced · Updated: 2026-10-03  
Publisher: TradeLabs AI (https://tradelabsai.com). Education, not financial advice.  
Cite as: TradeLabs Learn, "Data Versioning, Lineage and Schemas", https://learn.tradelabsai.com/programming/data-versioning/

Six months after a promising backtest, can you reproduce it exactly? Data vendors revise history, pipelines change, adjustment factors update with each dividend and code evolves. Without versioning, a strategy's results can drift for reasons nobody can trace, and it becomes impossible to tell whether a change in performance comes from the market or from the data. Data versioning means recording exactly which version of the data, code and parameters produced each result, so any result can be rebuilt and compared.

## What needs versioning

| Item | Why it changes | Tool |
|---|---|---|
| Code | Bug fixes and new features | git |
| Parameters and configuration | Tuning and experiments | git, config files |
| Raw data | Vendor corrections, new downloads | Snapshots, DVC, object storage versions |
| Processed data | Pipeline changes | Rebuild from versioned raw data and code |
| Environment | Library upgrades change results | Lock files, containers. See [Docker and Kubernetes](https://learn.tradelabsai.com/infrastructure/docker-and-kubernetes/) |
| Results | Each run's output | Experiment logs |

## Approaches

| Approach | How it works | Good for |
|---|---|---|
| Dated snapshots | Save each download in a folder named by date | Simple, small to medium datasets |
| Content hashes | Fingerprint each file; any change produces a new hash | Detecting silent changes |
| Append only tables | Never overwrite; add rows with a version or timestamp | Databases. See [Point-in-Time and Survivorship-Free Data](https://learn.tradelabsai.com/programming/point-in-time-data/) |
| Data version control tools | DVC, lakeFS or similar track large files alongside git | Larger research teams |
| Table formats with time travel | Formats such as Delta Lake or Apache Iceberg keep historical versions | Large data lakes |

## A simple workflow for individuals

1. **Keep raw downloads** in dated folders and never edit them.
2. **Compute a hash** (such as SHA 256) for each dataset used in a backtest.
3. **Commit code and configuration** to git before each important run.
4. **Log each run** with the git commit ID, data hashes, parameters and results.
5. **Pin library versions** in a requirements or lock file.

```python
import hashlib, json, subprocess

def file_hash(path):
    h = hashlib.sha256()
    with open(path, "rb") as f:
        for chunk in iter(lambda: f.read(1 << 20), b""):
            h.update(chunk)
    return h.hexdigest()[:16]

run_record = {
    "commit": subprocess.check_output(["git", "rev-parse", "HEAD"]).decode().strip(),
    "data": {"prices.parquet": file_hash("prices.parquet")},
    "params": {"fast": 20, "slow": 50},
}
print(json.dumps(run_record))
```

**Example: Tracing a result that changed**
A trader reruns a strategy and gets a Sharpe ratio of 0.9 instead of last quarter's 1.3. The run log shows the code commit is the same but the data hash differs. Comparing the two price snapshots reveals that the vendor corrected a block of bad prices in one stock during 2021; the old data contained a false spike that the strategy had "profited" from. The lower figure is the honest one. Without the stored hashes and snapshots, the trader might have spent days hunting for a code bug, or worse, trusted the old number. See [Cleaning Market Data](https://learn.tradelabsai.com/programming/cleaning-market-data/).

## Versioning and point in time data

Versioning answers "what did my dataset look like when I ran this?" Point in time data answers "what was known in the market on that date?" Both are needed: point in time data protects against look ahead bias, and versioning protects reproducibility. See [Backtest Reproducibility](https://learn.tradelabsai.com/research/backtest-reproducibility/).

## Storage considerations

Full copies of large datasets can become expensive. Options include storing only changed partitions, keeping compressed columnar files and pruning old snapshots that no result depends on. Cloud object storage with versioning enabled is a simple, low maintenance choice. See [Data Storage, Compression and Caching](https://learn.tradelabsai.com/programming/data-storage/).

## Frequently asked questions

### What is data versioning?

Recording which exact version of data was used for each analysis, so results can be reproduced and changes in data can be detected.

### Why do backtest results change when I rerun them?

Data revisions, changed adjustment factors, code edits and library upgrades can all change results; versioning each one shows which caused the change.

### What tools are used for data versioning?

git for code, dated snapshots and hashes for simple setups, and tools such as DVC, lakeFS, Delta Lake or Apache Iceberg for larger datasets.

Next, learn where and how to store market data efficiently in [Data Storage, Compression and Caching](https://learn.tradelabsai.com/programming/data-storage/).

## Continue learning

- Next lesson: [Data Storage, Compression and Caching](https://learn.tradelabsai.com/programming/data-storage/)
- Previous lesson: [Point-in-Time and Survivorship-Free Data](https://learn.tradelabsai.com/programming/point-in-time-data/)
- Related: [Point-in-Time and Survivorship-Free Data](https://learn.tradelabsai.com/programming/point-in-time-data/): Point in time data records what was known on each date, including restated figures and index changes. Learn why it matters and how to build point in time datasets.
- Related: [Backtest Reproducibility](https://learn.tradelabsai.com/research/backtest-reproducibility/): A reproducible backtest gives the same results every time from the same code and data. Learn version control, data snapshots, research logs and good habits.
- Related: [Data Pipelines and ETL](https://learn.tradelabsai.com/programming/data-pipelines-and-etl/): How trading data pipelines extract, transform and load market data reliably. Learn pipeline stages, scheduling, idempotent loads, validation checks and monitoring.
- Related: [Data Storage, Compression and Caching](https://learn.tradelabsai.com/programming/data-storage/): Compare ways to store market data: CSV, Parquet, HDF5, PostgreSQL, time series and columnar databases. Learn compression, partitioning and how to choose.
- Related: [The Trading Research Process](https://learn.tradelabsai.com/research/the-trading-research-process/): A disciplined research process turns ideas into tested strategies. Learn each step, from hypothesis and data to backtests, validation and paper trading.
