# NumPy and Pandas for Traders

> Learn the pandas and NumPy operations traders use most: loading price data, returns, rolling windows, resampling bars, joining assets and avoiding common traps.

Source: https://learn.tradelabsai.com/programming/numpy-and-pandas-for-traders/  
Track: Programming and Data · Level: Intermediate · Updated: 2026-10-03  
Publisher: TradeLabs AI (https://tradelabsai.com). Education, not financial advice.  
Cite as: TradeLabs Learn, "NumPy and Pandas for Traders", https://learn.tradelabsai.com/programming/numpy-and-pandas-for-traders/

NumPy and pandas are the two Python libraries behind almost every trading analysis. NumPy provides fast arrays and maths; pandas builds on it with labelled tables called DataFrames, which are ideal for time series such as prices indexed by date. Once you know a dozen core operations, you can calculate returns, indicators, rolling statistics and simple backtests in a few lines each, running hundreds of times faster than plain Python loops.

## The core objects

| Object | Library | Think of it as |
|---|---|---|
| ndarray | NumPy | A fast grid of numbers |
| Series | pandas | One column with an index, such as closing prices by date |
| DataFrame | pandas | A table of columns sharing an index, such as open, high, low, close and volume |
| DatetimeIndex | pandas | A time index that enables resampling and date slicing |

## Loading and inspecting data

```python
import pandas as pd
import numpy as np

df = pd.read_csv("btc_1h.csv", parse_dates=["time"], index_col="time")
df = df.sort_index()
print(df.head())
print(df.isna().sum())        # missing values per column
print(df.index.is_unique)     # duplicate timestamps?
```

Always sort by time and check for missing values and duplicate timestamps before doing anything else. See [Cleaning Market Data](https://learn.tradelabsai.com/programming/cleaning-market-data/).

## The operations traders use most

| Task | pandas code |
|---|---|
| Simple returns | `df["Close"].pct_change()` |
| Log returns | `np.log(df["Close"]).diff()` |
| Moving average | `df["Close"].rolling(20).mean()` |
| Rolling volatility | `df["ret"].rolling(20).std() * np.sqrt(252)` |
| Exponential average | `df["Close"].ewm(span=20).mean()` |
| Previous value | `df["Close"].shift(1)` |
| Cumulative growth | `(1 + df["ret"]).cumprod()` |
| Running peak and drawdown | `eq / eq.cummax() - 1` |
| Date slice | `df.loc["2025-01":"2025-06"]` |

See [Rolling and Expanding Windows](https://learn.tradelabsai.com/math/rolling-and-expanding-windows/) and [Measuring Returns and CAGR](https://learn.tradelabsai.com/portfolio/measuring-returns-and-cagr/).

## Resampling bars

Turning 1 minute bars into 1 hour bars needs a different rule for each column:

```python
hourly = df.resample("1h").agg({
    "Open": "first", "High": "max", "Low": "min",
    "Close": "last", "Volume": "sum",
}).dropna()
```

Be careful with labels: by default pandas labels each hourly bar by its start time. If your strategy treats the label as the time the bar is complete, you introduce look ahead. See [Tick Data and OHLCV Data](https://learn.tradelabsai.com/programming/tick-data-and-ohlcv-data/) and [Timestamps, Time Zones and Daylight Saving](https://learn.tradelabsai.com/programming/timestamps-and-time-zones/).

## Combining several assets

```python
closes = pd.concat({"SPY": spy["Close"], "TLT": tlt["Close"]}, axis=1)
rets = closes.pct_change().dropna()
print(rets.corr())
```

`concat` aligns on dates automatically. Rows where one asset is missing appear as NaN; decide deliberately whether to drop them or fill them. Forward filling a price is usually acceptable for a holiday; forward filling a return is not. See [Covariance and Correlation](https://learn.tradelabsai.com/math/covariance-and-correlation/).

**Example: Annualised volatility by hand**
Daily returns for a stock over 20 days have a standard deviation of 1.5%. Annualising with the usual 252 trading days gives 1.5% times the square root of 252, about 1.5% times 15.87, or roughly 23.8%. In pandas this is one line: `rets.rolling(20).std() * np.sqrt(252)`. For crypto, which trades every day, many analysts use 365 instead, which would give about 28.7%. See [Historical and Realized Volatility](https://learn.tradelabsai.com/volatility/historical-volatility/).

## Why vectorized code is faster

A loop in Python processes one number at a time through the interpreter. NumPy and pandas operations hand the whole array to optimised compiled code. On a million rows, a rolling mean in pandas typically finishes in milliseconds while an equivalent Python loop can take seconds. See [Event-Driven vs Vectorized Backtesting](https://learn.tradelabsai.com/research/vectorized-backtesting/).

## Common traps

1. **Chained assignment** such as `df[df.a > 0]["b"] = 1`, which may silently not change `df`. Use `df.loc[df.a > 0, "b"] = 1`.
2. **Time zone mixing** between naive and aware timestamps.
3. **Forgetting `shift`** when turning a signal into a position.
4. **Dropping NaNs too early,** which misaligns series.
5. **Using floats for exact money** in accounting code; fine for research, risky for ledgers.

## A practice routine

The fastest way to learn pandas is to answer real questions with it. Download a few years of daily prices for one stock and one index. Then calculate daily returns, the largest one day gain and loss, the average volume by weekday, the 20 day rolling volatility and the correlation between the two assets for each calendar year. Next, resample the daily data into weekly bars and check that the weekly highs equal the highest daily high in each week. Every answer you can verify by hand on a spreadsheet builds trust in your code. Once these feel natural, move on to a full backtest with costs in [Backtesting Methodology](https://learn.tradelabsai.com/research/backtesting-methodology/), and keep each analysis in a script under version control so you can rerun it on new data.

## Frequently asked questions

### What is pandas used for in trading?

Loading, cleaning and transforming price and trade data, calculating returns and indicators, resampling bars and running simple backtests.

### Do I need NumPy if I use pandas?

pandas uses NumPy underneath, and you will use NumPy functions such as `np.log` and `np.sqrt` alongside pandas regularly.

### How do I calculate returns in pandas?

Use `pct_change()` for simple returns or `np.log(prices).diff()` for log returns.

Next, learn to chart your results in [Plotting Market Data with Matplotlib](https://learn.tradelabsai.com/programming/matplotlib/).

## Continue learning

- Next lesson: [Plotting Market Data with Matplotlib](https://learn.tradelabsai.com/programming/matplotlib/)
- Previous lesson: [Python for Trading](https://learn.tradelabsai.com/programming/python-for-trading/)
- Related: [Python for Trading](https://learn.tradelabsai.com/programming/python-for-trading/): Why Python is the most popular language for trading research and bots, which libraries matter, how to set up a project and a first script that tests a simple rule.
- Related: [Plotting Market Data with Matplotlib](https://learn.tradelabsai.com/programming/matplotlib/): Use matplotlib to plot prices, indicators, equity curves, drawdowns and return histograms, with clear examples and tips for honest, readable trading charts.
- Related: [Event-Driven vs Vectorized Backtesting](https://learn.tradelabsai.com/research/vectorized-backtesting/): Vectorised backtests compute signals and returns for all dates at once using arrays. Learn how they work, a pandas example, their speed advantages and their traps.
- Related: [Rolling and Expanding Windows](https://learn.tradelabsai.com/math/rolling-and-expanding-windows/): Rolling windows use a fixed recent period; expanding windows use all data so far. Learn when to use each, window length trade offs and how to avoid look ahead bias.
- Related: [Tick Data and OHLCV Data](https://learn.tradelabsai.com/programming/tick-data-and-ohlcv-data/): Tick data records every trade or quote; OHLCV bars summarise them by time. Learn how bars are built, other bar types, storage costs and which data a strategy needs.
