TradeLabs AILearn

Maximum Likelihood

Maximum likelihood estimation finds the model parameters that make observed data most probable. Learn the idea, simple examples, its use in GARCH and its limits.

Intermediate3 min readUpdated 3 Oct 2026
Markdown
Lesson 20 of 46

Maximum likelihood estimation (MLE) is a general method for fitting statistical models to data. The idea is simple: among all possible parameter values, choose the ones that make the data you actually observed most probable. MLE is used to estimate volatility models like GARCH, fit return distributions, calibrate option pricing models and estimate the parameters of many machine learning models. Understanding the idea helps traders interpret model outputs and recognise their limits.

The idea#

Suppose you have data and a model with unknown parameters. The likelihood function measures how probable the data would be for each possible parameter value.

L(θ) = P(data | θ) = Π f(x_i | θ)
log L(θ) = Σ log f(x_i | θ)

MLE chooses the θ that maximises L(θ). In practice, we maximise the log likelihood, which turns products into sums and is easier to work with.

A simple example: a coin or a win rate#

Normal distribution#

For data assumed to be normal, MLE gives:

μ̂ = sample mean
σ̂² = Σ (x - μ̂)² / n

Note the n rather than n minus 1: the MLE of variance is slightly biased in small samples, which is why the sample variance formula uses n minus 1. See Variance and Standard Deviation.

MLE in finance#

ApplicationWhat is estimatedLesson
GARCH volatility modelsParameters controlling how volatility reacts and persistsGARCH
Fat tailed distributionsDegrees of freedom of a Student t distributionStudent's t-Distribution
ARIMA modelsTime series coefficientsARIMA
Logistic regressionProbability of an event, such as default or a price riseRegression and Classification Models
Regime switching modelsProbabilities and parameters of hidden regimesStructural Breaks and Regime Changes
Option model calibrationParameters of stochastic volatility models (often with related methods)Stochastic Volatility and the Heston Model
  • General: works for almost any model with a defined probability distribution.
  • Efficient: in large samples, MLE estimates have the smallest possible variance under standard conditions.
  • Standard errors: the curvature of the log likelihood gives estimates of uncertainty.
  • Model comparison: likelihood based criteria such as AIC and BIC compare models while penalising complexity.
AIC = 2k - 2 log L
BIC = k log(n) - 2 log L

where k is the number of parameters and n the number of observations. Lower values are better.

Limits and pitfalls#

IssueExplanation
Wrong modelMLE finds the best parameters for the chosen model, even if the model is wrong
Small samplesEstimates can be biased and unstable
Local maximaComplex likelihoods can have several peaks; optimisers may get stuck. See Optimization
OverfittingMore parameters always increase the likelihood; use AIC, BIC or out of sample tests
Fat tails and outliersNormal based MLE is sensitive to extreme values
Changing marketsParameters estimated on the past may not hold. See Structural Breaks and Regime Changes

MLE vs Bayesian estimation#

MLE uses only the data. Bayesian estimation combines the likelihood with a prior distribution, which can stabilise estimates when data is limited. With lots of data, the two usually agree. See Bayesian Statistics.

Frequently asked questions#

What is maximum likelihood estimation?#

A method that chooses model parameters that make the observed data most probable under the model.

Where is maximum likelihood used in trading?#

In estimating volatility models like GARCH, fitting return distributions, time series models, logistic regressions and calibrating some pricing models.

What are the limits of maximum likelihood?#

It assumes the model is correct, can be unstable in small samples, may find local rather than global maxima and can overfit with many parameters.

Next, learn the Bayesian approach in Bayesian Statistics.

Check your understanding

3 quick questions on this lesson. Get them all right to finish it.

Turn on JavaScript to take the quiz.

Finished this lesson?Sign in to save your progress across devices.
Next lessonBayesian StatisticsBayesian statistics combines prior beliefs with data to estimate uncertain quantities. Learn priors and posteriors, shrinkage, credible intervals and trading uses.

Mentioned in