# Data Storage, Compression and Caching

> Compare ways to store market data: CSV, Parquet, HDF5, PostgreSQL, time series and columnar databases. Learn compression, partitioning and how to choose.

Source: https://learn.tradelabsai.com/programming/data-storage/  
Track: Programming and Data · Level: Advanced · Updated: 2026-10-03  
Publisher: TradeLabs AI (https://tradelabsai.com). Education, not financial advice.  
Cite as: TradeLabs Learn, "Data Storage, Compression and Caching", https://learn.tradelabsai.com/programming/data-storage/

Where and how you store market data decides how fast your research runs and how much it costs. A few years of daily bars fit comfortably in a spreadsheet; a few years of tick data for an active market can reach terabytes. The right storage depends on data size, how often you write, how you query and whether the data feeds research, live trading or both. Most traders end up with a combination: compact files for research, a database for orders and live state, and an in memory cache for the newest prices.

## File formats

| Format | Type | Strengths | Weaknesses |
|---|---|---|---|
| CSV | Text, row based | Readable anywhere | Large, slow, no types, timestamp ambiguity |
| Parquet | Binary, columnar | Compressed, fast column reads, typed, widely supported | Not ideal for frequent small appends |
| Feather / Arrow | Binary, columnar | Very fast to read into pandas | Less compression than Parquet |
| HDF5 | Binary | Fast for large arrays | Fussier tooling, file locking issues |

For research datasets, Parquet is the usual default today. pandas reads and writes it with `read_parquet` and `to_parquet`.

## Why columnar storage is fast

Row formats store each record together: time, open, high, low, close, volume, then the next row. Columnar formats store each column together. If a query needs only the close prices, a columnar file reads just that column. Similar values stored next to each other also compress very well, often several times smaller than CSV.

**Example: CSV versus Parquet for minute bars**
A year of one minute bars for 500 stocks is about 49 million rows (500 symbols times 390 minutes times about 252 days). As CSV, at roughly 60 bytes per row, that is about 3 GB. As compressed Parquet, the same data commonly shrinks to a fraction of that size, and loading only the close column reads a small part of the file. Actual ratios depend on the data, so measure on your own files, but the difference is usually large enough to change how fast research feels. See [Tick Data and OHLCV Data](https://learn.tradelabsai.com/programming/tick-data-and-ohlcv-data/).

## Databases

| Type | Examples | Best for |
|---|---|---|
| Relational | PostgreSQL, MySQL, SQLite | Orders, fills, accounts, reference data, moderate bar data. See [SQL for Trading Data](https://learn.tradelabsai.com/programming/sql-for-trading-data/) |
| Time series extension | TimescaleDB on PostgreSQL | Large bar and tick tables with SQL |
| Columnar analytical | ClickHouse, DuckDB | Fast analytics over billions of rows |
| Specialist time series | kdb+, QuestDB, InfluxDB | Very high write rates; kdb+ is common in banks and trading firms |
| In memory | Redis | Latest prices, live state, caches. See [Redis and PostgreSQL for Trading](https://learn.tradelabsai.com/infrastructure/redis-and-postgresql-for-trading/) |

DuckDB deserves a mention for individual researchers: it runs inside Python with no server and can query Parquet files directly with SQL.

## Partitioning

Split large datasets into partitions, usually by date and sometimes by symbol, such as one Parquet file per day or per month per symbol. Queries for a date range then read only the relevant partitions, and new days are added without rewriting old files. See [Database Design for Market Data](https://learn.tradelabsai.com/programming/database-design-for-market-data/).

## Hot, warm and cold data

| Tier | Data | Storage |
|---|---|---|
| Hot | Today's live prices, open orders | Memory, Redis |
| Warm | Recent months used in daily research | Local SSD, database |
| Cold | Older history, raw archives | Object storage, compressed files |

## Choosing a setup

| Situation | Suggested starting point |
|---|---|
| Daily bars, a few hundred symbols | Parquet files or SQLite |
| Minute bars, thousands of symbols | Partitioned Parquet with DuckDB, or TimescaleDB |
| Tick data research | Partitioned Parquet or a columnar database |
| Live bot state, orders and fills | PostgreSQL |
| Latest prices shared between processes | Redis |

## Backups and integrity

Keep raw data separate and backed up, verify files with checksums and test restores. See [Failover, Backups and Disaster Recovery](https://learn.tradelabsai.com/algo-trading/disaster-recovery/) and [Data Versioning, Lineage and Schemas](https://learn.tradelabsai.com/programming/data-versioning/).

## Frequently asked questions

### What is the best format to store stock data?

For research, compressed columnar Parquet files are a strong default; for orders and live state, a relational database such as PostgreSQL.

### Is CSV good for market data?

It is fine for small datasets and sharing, but it is large, slow and loses type information, so larger datasets are better stored as Parquet or in a database.

### What database do trading firms use?

Many use relational databases for orders and accounts and specialist time series or columnar databases, such as kdb+ or ClickHouse, for market data.

Next, learn how order book data is delivered in [Order Book Feeds: Snapshots and Incremental Updates](https://learn.tradelabsai.com/programming/order-book-feeds/).

## Continue learning

- Next lesson: [Order Book Feeds: Snapshots and Incremental Updates](https://learn.tradelabsai.com/programming/order-book-feeds/)
- Previous lesson: [Data Versioning, Lineage and Schemas](https://learn.tradelabsai.com/programming/data-versioning/)
- Related: [Data Versioning, Lineage and Schemas](https://learn.tradelabsai.com/programming/data-versioning/): Data versioning tracks exactly which data, code and settings produced each backtest. Learn snapshots, hashes, tools like git and DVC, and a simple workflow.
- Related: [Database Design for Market Data](https://learn.tradelabsai.com/programming/database-design-for-market-data/): How to design database tables for bars, ticks, symbols, orders and fills. Learn keys, indexes, partitioning, data types and how to avoid common design mistakes.
- Related: [Redis and PostgreSQL for Trading](https://learn.tradelabsai.com/infrastructure/redis-and-postgresql-for-trading/): How Redis and PostgreSQL work together in trading systems: Redis for live prices, caches and state, PostgreSQL for durable orders, fills and history.
- Related: [Tick Data and OHLCV Data](https://learn.tradelabsai.com/programming/tick-data-and-ohlcv-data/): Tick data records every trade or quote; OHLCV bars summarise them by time. Learn how bars are built, other bar types, storage costs and which data a strategy needs.
- Related: [SQL for Trading Data](https://learn.tradelabsai.com/programming/sql-for-trading-data/): Learn the SQL queries traders use most: filtering bars, aggregating trades into candles, joining fills to orders, window functions for returns and daily P&L.
