TradeLabs AILearn

Data Versioning, Lineage and Schemas

Data versioning tracks exactly which data, code and settings produced each backtest. Learn snapshots, hashes, tools like git and DVC, and a simple workflow.

Advanced3 min readUpdated 3 Oct 2026
Markdown
Lesson 21 of 27

Six months after a promising backtest, can you reproduce it exactly? Data vendors revise history, pipelines change, adjustment factors update with each dividend and code evolves. Without versioning, a strategy's results can drift for reasons nobody can trace, and it becomes impossible to tell whether a change in performance comes from the market or from the data. Data versioning means recording exactly which version of the data, code and parameters produced each result, so any result can be rebuilt and compared.

What needs versioning#

ItemWhy it changesTool
CodeBug fixes and new featuresgit
Parameters and configurationTuning and experimentsgit, config files
Raw dataVendor corrections, new downloadsSnapshots, DVC, object storage versions
Processed dataPipeline changesRebuild from versioned raw data and code
EnvironmentLibrary upgrades change resultsLock files, containers. See Docker and Kubernetes
ResultsEach run's outputExperiment logs

Approaches#

ApproachHow it worksGood for
Dated snapshotsSave each download in a folder named by dateSimple, small to medium datasets
Content hashesFingerprint each file; any change produces a new hashDetecting silent changes
Append only tablesNever overwrite; add rows with a version or timestampDatabases. See Point-in-Time and Survivorship-Free Data
Data version control toolsDVC, lakeFS or similar track large files alongside gitLarger research teams
Table formats with time travelFormats such as Delta Lake or Apache Iceberg keep historical versionsLarge data lakes

A simple workflow for individuals#

  1. Keep raw downloads in dated folders and never edit them.
  2. Compute a hash (such as SHA 256) for each dataset used in a backtest.
  3. Commit code and configuration to git before each important run.
  4. Log each run with the git commit ID, data hashes, parameters and results.
  5. Pin library versions in a requirements or lock file.
import hashlib, json, subprocess

def file_hash(path):
    h = hashlib.sha256()
    with open(path, "rb") as f:
        for chunk in iter(lambda: f.read(1 << 20), b""):
            h.update(chunk)
    return h.hexdigest()[:16]

run_record = {
    "commit": subprocess.check_output(["git", "rev-parse", "HEAD"]).decode().strip(),
    "data": {"prices.parquet": file_hash("prices.parquet")},
    "params": {"fast": 20, "slow": 50},
}
print(json.dumps(run_record))

Versioning and point in time data#

Versioning answers "what did my dataset look like when I ran this?" Point in time data answers "what was known in the market on that date?" Both are needed: point in time data protects against look ahead bias, and versioning protects reproducibility. See Backtest Reproducibility.

Storage considerations#

Full copies of large datasets can become expensive. Options include storing only changed partitions, keeping compressed columnar files and pruning old snapshots that no result depends on. Cloud object storage with versioning enabled is a simple, low maintenance choice. See Data Storage, Compression and Caching.

Frequently asked questions#

What is data versioning?#

Recording which exact version of data was used for each analysis, so results can be reproduced and changes in data can be detected.

Why do backtest results change when I rerun them?#

Data revisions, changed adjustment factors, code edits and library upgrades can all change results; versioning each one shows which caused the change.

What tools are used for data versioning?#

git for code, dated snapshots and hashes for simple setups, and tools such as DVC, lakeFS, Delta Lake or Apache Iceberg for larger datasets.

Next, learn where and how to store market data efficiently in Data Storage, Compression and Caching.

Check your understanding

3 quick questions on this lesson. Get them all right to finish it.

Turn on JavaScript to take the quiz.

Finished this lesson?Sign in to save your progress across devices.
Next lessonData Storage, Compression and CachingCompare ways to store market data: CSV, Parquet, HDF5, PostgreSQL, time series and columnar databases. Learn compression, partitioning and how to choose.

Mentioned in