Data Storage, Compression and Caching
Compare ways to store market data: CSV, Parquet, HDF5, PostgreSQL, time series and columnar databases. Learn compression, partitioning and how to choose.
Where and how you store market data decides how fast your research runs and how much it costs. A few years of daily bars fit comfortably in a spreadsheet; a few years of tick data for an active market can reach terabytes. The right storage depends on data size, how often you write, how you query and whether the data feeds research, live trading or both. Most traders end up with a combination: compact files for research, a database for orders and live state, and an in memory cache for the newest prices.
File formats#
| Format | Type | Strengths | Weaknesses |
|---|---|---|---|
| CSV | Text, row based | Readable anywhere | Large, slow, no types, timestamp ambiguity |
| Parquet | Binary, columnar | Compressed, fast column reads, typed, widely supported | Not ideal for frequent small appends |
| Feather / Arrow | Binary, columnar | Very fast to read into pandas | Less compression than Parquet |
| HDF5 | Binary | Fast for large arrays | Fussier tooling, file locking issues |
For research datasets, Parquet is the usual default today. pandas reads and writes it with read_parquet and to_parquet.
Why columnar storage is fast#
Row formats store each record together: time, open, high, low, close, volume, then the next row. Columnar formats store each column together. If a query needs only the close prices, a columnar file reads just that column. Similar values stored next to each other also compress very well, often several times smaller than CSV.
Databases#
| Type | Examples | Best for |
|---|---|---|
| Relational | PostgreSQL, MySQL, SQLite | Orders, fills, accounts, reference data, moderate bar data. See SQL for Trading Data |
| Time series extension | TimescaleDB on PostgreSQL | Large bar and tick tables with SQL |
| Columnar analytical | ClickHouse, DuckDB | Fast analytics over billions of rows |
| Specialist time series | kdb+, QuestDB, InfluxDB | Very high write rates; kdb+ is common in banks and trading firms |
| In memory | Redis | Latest prices, live state, caches. See Redis and PostgreSQL for Trading |
DuckDB deserves a mention for individual researchers: it runs inside Python with no server and can query Parquet files directly with SQL.
Partitioning#
Split large datasets into partitions, usually by date and sometimes by symbol, such as one Parquet file per day or per month per symbol. Queries for a date range then read only the relevant partitions, and new days are added without rewriting old files. See Database Design for Market Data.
Hot, warm and cold data#
| Tier | Data | Storage |
|---|---|---|
| Hot | Today's live prices, open orders | Memory, Redis |
| Warm | Recent months used in daily research | Local SSD, database |
| Cold | Older history, raw archives | Object storage, compressed files |
Choosing a setup#
| Situation | Suggested starting point |
|---|---|
| Daily bars, a few hundred symbols | Parquet files or SQLite |
| Minute bars, thousands of symbols | Partitioned Parquet with DuckDB, or TimescaleDB |
| Tick data research | Partitioned Parquet or a columnar database |
| Live bot state, orders and fills | PostgreSQL |
| Latest prices shared between processes | Redis |
Backups and integrity#
Keep raw data separate and backed up, verify files with checksums and test restores. See Failover, Backups and Disaster Recovery and Data Versioning, Lineage and Schemas.
Frequently asked questions#
What is the best format to store stock data?#
For research, compressed columnar Parquet files are a strong default; for orders and live state, a relational database such as PostgreSQL.
Is CSV good for market data?#
It is fine for small datasets and sharing, but it is large, slow and loses type information, so larger datasets are better stored as Parquet or in a database.
What database do trading firms use?#
Many use relational databases for orders and accounts and specialist time series or columnar databases, such as kdb+ or ClickHouse, for market data.
Next, learn how order book data is delivered in Order Book Feeds: Snapshots and Incremental Updates.
3 quick questions on this lesson. Get them all right to finish it.
Turn on JavaScript to take the quiz.
Mentioned in
- Data Pipelines and ETLProgramming and Data
- Historical Data for BacktestingResearch and Backtesting
- Quant Developer and Trading EngineerThe Trading Industry