Online Learning
Online learning updates a model with each new observation instead of retraining in batches. Learn how it works, forgetting factors, drift detection and its risks.
Most machine learning models are trained in batches: collect data, train, deploy, then retrain later. Online learning instead updates the model continuously, a little with each new observation. Because markets change, a model that adapts as data arrives sounds ideal. In practice, online learning is a trade off: adapt too slowly and the model lags behind changing conditions; adapt too quickly and it chases noise. Done well, it keeps simple models current with little computing cost.
Batch versus online learning#
| Batch learning | Online learning | |
|---|---|---|
| Updates | Periodic retrains on large datasets | Small update after each observation or mini batch |
| Adaptation | Stepwise | Continuous |
| Computing | Heavy at retrain times | Light and constant |
| Memory | Needs stored history | Can work with only the current model state |
| Risk | Stale between retrains | Drift toward noise, harder to audit |
Simple online methods#
| Method | How it adapts |
|---|---|
| Exponentially weighted averages | Recent data gets more weight. See Exponential Moving Average (EMA) |
| Recursive least squares | Updates regression coefficients with each new point, often with a forgetting factor |
| Stochastic gradient descent | Takes a small learning step on each new example |
| Kalman filter | Tracks changing parameters, such as a hedge ratio, with an explicit model of noise |
| Online ensembles | Reweights several models based on recent performance |
The forgetting factor#
Many online methods use a forgetting factor that controls how quickly old data loses influence. An exponential weight with factor lambda gives an effective memory of roughly 1 divided by (1 minus lambda) observations.
| Lambda | Effective memory |
|---|---|
| 0.90 | About 10 observations |
| 0.99 | About 100 observations |
| 0.999 | About 1,000 observations |
Shorter memory adapts faster but is noisier.
Detecting drift#
Rather than adapting constantly, some systems monitor for concept drift: a change in the relationship between features and targets. When a drift detector signals a significant change in prediction errors, the system retrains or switches models. This combines stability in calm periods with quick response to real shifts. See Structural Breaks and Regime Changes.
Risks of online learning#
- Chasing noise: fast adaptation fits random fluctuations.
- Feedback loops: if the model's own trades affect prices, it may learn from its own impact.
- Silent degradation: a model that keeps changing is harder to monitor and audit. See Operational and Model Risk.
- Bad data poisoning: a burst of bad ticks can push the model off course. See Cleaning Market Data.
- Validation complexity: you must simulate updates step by step to test it honestly.
Safeguards#
- Bound parameter changes per update.
- Log every model state so behaviour can be reconstructed. See Logging, Audit Trails and Incident Response.
- Compare with a frozen benchmark model to detect degradation.
- Validate the full online process with walk forward replay. See Walk-Forward Validation and Preventing Overfitting.
- Filter inputs for bad data before updating.
Frequently asked questions#
What is online learning in trading?#
A machine learning approach where the model updates continuously with each new data point instead of being retrained in periodic batches.
Is online learning better than retraining?#
Not always. It adapts faster but can chase noise; scheduled retraining is simpler to validate and monitor. Many systems combine both.
What is a forgetting factor?#
A setting that controls how quickly old observations lose weight in an adaptive model; values closer to 1 give longer memory.
Next, learn about agents that learn by trial and error in Reinforcement Learning.
3 quick questions on this lesson. Get them all right to finish it.
Turn on JavaScript to take the quiz.