# Online Learning

> Online learning updates a model with each new observation instead of retraining in batches. Learn how it works, forgetting factors, drift detection and its risks.

Source: https://learn.tradelabsai.com/machine-learning/online-learning/  
Track: Machine Learning · Level: Advanced · Updated: 2026-10-03  
Publisher: TradeLabs AI (https://tradelabsai.com). Education, not financial advice.  
Cite as: TradeLabs Learn, "Online Learning", https://learn.tradelabsai.com/machine-learning/online-learning/

Most machine learning models are trained in batches: collect data, train, deploy, then retrain later. Online learning instead updates the model continuously, a little with each new observation. Because markets change, a model that adapts as data arrives sounds ideal. In practice, online learning is a trade off: adapt too slowly and the model lags behind changing conditions; adapt too quickly and it chases noise. Done well, it keeps simple models current with little computing cost.

## Batch versus online learning

| | Batch learning | Online learning |
|---|---|---|
| Updates | Periodic retrains on large datasets | Small update after each observation or mini batch |
| Adaptation | Stepwise | Continuous |
| Computing | Heavy at retrain times | Light and constant |
| Memory | Needs stored history | Can work with only the current model state |
| Risk | Stale between retrains | Drift toward noise, harder to audit |

## Simple online methods

| Method | How it adapts |
|---|---|
| Exponentially weighted averages | Recent data gets more weight. See [Exponential Moving Average (EMA)](https://learn.tradelabsai.com/indicators/exponential-moving-average/) |
| Recursive least squares | Updates regression coefficients with each new point, often with a forgetting factor |
| Stochastic gradient descent | Takes a small learning step on each new example |
| Kalman filter | Tracks changing parameters, such as a hedge ratio, with an explicit model of noise |
| Online ensembles | Reweights several models based on recent performance |

## The forgetting factor

Many online methods use a forgetting factor that controls how quickly old data loses influence. An exponential weight with factor lambda gives an effective memory of roughly 1 divided by (1 minus lambda) observations.

| Lambda | Effective memory |
|---|---|
| 0.90 | About 10 observations |
| 0.99 | About 100 observations |
| 0.999 | About 1,000 observations |

Shorter memory adapts faster but is noisier.

**Example: Adaptive hedge ratio for a pair**
A pairs trader tracks the hedge ratio between two related stocks. A fixed ratio estimated on two years of data is 1.20 shares of stock B per share of stock A. A Kalman filter updates the ratio daily as new prices arrive. Over six months, after one company changes its business mix, the filter's estimate drifts to 1.05. The trading spread built with the adaptive ratio stays mean reverting, while the spread built with the fixed 1.20 ratio trends steadily, producing losing trades. Online estimation kept the model in line with the changed relationship. See [Pairs Trading](https://learn.tradelabsai.com/strategies/pairs-trading/) and [Cointegration](https://learn.tradelabsai.com/math/cointegration/).

## Detecting drift

Rather than adapting constantly, some systems monitor for concept drift: a change in the relationship between features and targets. When a drift detector signals a significant change in prediction errors, the system retrains or switches models. This combines stability in calm periods with quick response to real shifts. See [Structural Breaks and Regime Changes](https://learn.tradelabsai.com/math/regime-changes/).

## Risks of online learning

- **Chasing noise:** fast adaptation fits random fluctuations.
- **Feedback loops:** if the model's own trades affect prices, it may learn from its own impact.
- **Silent degradation:** a model that keeps changing is harder to monitor and audit. See [Operational and Model Risk](https://learn.tradelabsai.com/portfolio/operational-and-model-risk/).
- **Bad data poisoning:** a burst of bad ticks can push the model off course. See [Cleaning Market Data](https://learn.tradelabsai.com/programming/cleaning-market-data/).
- **Validation complexity:** you must simulate updates step by step to test it honestly.

## Safeguards

1. **Bound parameter changes** per update.
2. **Log every model state** so behaviour can be reconstructed. See [Logging, Audit Trails and Incident Response](https://learn.tradelabsai.com/algo-trading/audit-trails/).
3. **Compare with a frozen benchmark model** to detect degradation.
4. **Validate the full online process** with walk forward replay. See [Walk-Forward Validation and Preventing Overfitting](https://learn.tradelabsai.com/machine-learning/walk-forward-validation/).
5. **Filter inputs** for bad data before updating.

## Frequently asked questions

### What is online learning in trading?

A machine learning approach where the model updates continuously with each new data point instead of being retrained in periodic batches.

### Is online learning better than retraining?

Not always. It adapts faster but can chase noise; scheduled retraining is simpler to validate and monitor. Many systems combine both.

### What is a forgetting factor?

A setting that controls how quickly old observations lose weight in an adaptive model; values closer to 1 give longer memory.

Next, learn about agents that learn by trial and error in [Reinforcement Learning](https://learn.tradelabsai.com/machine-learning/reinforcement-learning/).

## Continue learning

- Next lesson: [Reinforcement Learning](https://learn.tradelabsai.com/machine-learning/reinforcement-learning/)
- Previous lesson: [Walk-Forward Validation and Preventing Overfitting](https://learn.tradelabsai.com/machine-learning/walk-forward-validation/)
- Related: [Walk-Forward Validation and Preventing Overfitting](https://learn.tradelabsai.com/machine-learning/walk-forward-validation/): Walk forward validation retrains a model on a rolling or expanding window and tests it on the next period, just as it would be used live. Learn setup and choices.
- Related: [Structural Breaks and Regime Changes](https://learn.tradelabsai.com/math/regime-changes/): Markets switch between regimes such as calm and turbulent, or trending and ranging. Learn how to detect regimes, the models used and how to adapt strategies.
- Related: [Stationarity, Differencing and Unit Roots](https://learn.tradelabsai.com/math/stationarity/): A stationary series has stable statistical properties over time. Learn why prices are non stationary, how to test with ADF and KPSS and how to make data stationary.
- Related: [Signal and Alpha Decay](https://learn.tradelabsai.com/research/signal-and-alpha-decay/): Signal decay is how fast a signal's predictive power fades; alpha decay is how edges shrink over years. Learn both, the evidence and how traders adapt.
- Related: [Reinforcement Learning](https://learn.tradelabsai.com/machine-learning/reinforcement-learning/): Reinforcement learning trains agents to act by rewarding good outcomes. Learn how it applies to trading and execution, how rewards are designed and why it is hard.
