TradeLabs AILearn

Online Learning

Online learning updates a model with each new observation instead of retraining in batches. Learn how it works, forgetting factors, drift detection and its risks.

Advanced3 min readUpdated 3 Oct 2026
Markdown
Lesson 9 of 10

Most machine learning models are trained in batches: collect data, train, deploy, then retrain later. Online learning instead updates the model continuously, a little with each new observation. Because markets change, a model that adapts as data arrives sounds ideal. In practice, online learning is a trade off: adapt too slowly and the model lags behind changing conditions; adapt too quickly and it chases noise. Done well, it keeps simple models current with little computing cost.

Batch versus online learning#

Batch learningOnline learning
UpdatesPeriodic retrains on large datasetsSmall update after each observation or mini batch
AdaptationStepwiseContinuous
ComputingHeavy at retrain timesLight and constant
MemoryNeeds stored historyCan work with only the current model state
RiskStale between retrainsDrift toward noise, harder to audit

Simple online methods#

MethodHow it adapts
Exponentially weighted averagesRecent data gets more weight. See Exponential Moving Average (EMA)
Recursive least squaresUpdates regression coefficients with each new point, often with a forgetting factor
Stochastic gradient descentTakes a small learning step on each new example
Kalman filterTracks changing parameters, such as a hedge ratio, with an explicit model of noise
Online ensemblesReweights several models based on recent performance

The forgetting factor#

Many online methods use a forgetting factor that controls how quickly old data loses influence. An exponential weight with factor lambda gives an effective memory of roughly 1 divided by (1 minus lambda) observations.

LambdaEffective memory
0.90About 10 observations
0.99About 100 observations
0.999About 1,000 observations

Shorter memory adapts faster but is noisier.

Detecting drift#

Rather than adapting constantly, some systems monitor for concept drift: a change in the relationship between features and targets. When a drift detector signals a significant change in prediction errors, the system retrains or switches models. This combines stability in calm periods with quick response to real shifts. See Structural Breaks and Regime Changes.

Risks of online learning#

  • Chasing noise: fast adaptation fits random fluctuations.
  • Feedback loops: if the model's own trades affect prices, it may learn from its own impact.
  • Silent degradation: a model that keeps changing is harder to monitor and audit. See Operational and Model Risk.
  • Bad data poisoning: a burst of bad ticks can push the model off course. See Cleaning Market Data.
  • Validation complexity: you must simulate updates step by step to test it honestly.

Safeguards#

  1. Bound parameter changes per update.
  2. Log every model state so behaviour can be reconstructed. See Logging, Audit Trails and Incident Response.
  3. Compare with a frozen benchmark model to detect degradation.
  4. Validate the full online process with walk forward replay. See Walk-Forward Validation and Preventing Overfitting.
  5. Filter inputs for bad data before updating.

Frequently asked questions#

What is online learning in trading?#

A machine learning approach where the model updates continuously with each new data point instead of being retrained in periodic batches.

Is online learning better than retraining?#

Not always. It adapts faster but can chase noise; scheduled retraining is simpler to validate and monitor. Many systems combine both.

What is a forgetting factor?#

A setting that controls how quickly old observations lose weight in an adaptive model; values closer to 1 give longer memory.

Next, learn about agents that learn by trial and error in Reinforcement Learning.

Check your understanding

3 quick questions on this lesson. Get them all right to finish it.

Turn on JavaScript to take the quiz.

Finished this lesson?Sign in to save your progress across devices.
Next lessonReinforcement LearningReinforcement learning trains agents to act by rewarding good outcomes. Learn how it applies to trading and execution, how rewards are designed and why it is hard.