# Maximum Likelihood

> Maximum likelihood estimation finds the model parameters that make observed data most probable. Learn the idea, simple examples, its use in GARCH and its limits.

Source: https://learn.tradelabsai.com/math/maximum-likelihood/  
Track: Math and Statistics · Level: Intermediate · Updated: 2026-10-03  
Publisher: TradeLabs AI (https://tradelabsai.com). Education, not financial advice.  
Cite as: TradeLabs Learn, "Maximum Likelihood", https://learn.tradelabsai.com/math/maximum-likelihood/

Maximum likelihood estimation (MLE) is a general method for fitting statistical models to data. The idea is simple: among all possible parameter values, choose the ones that make the data you actually observed most probable. MLE is used to estimate volatility models like GARCH, fit return distributions, calibrate option pricing models and estimate the parameters of many machine learning models. Understanding the idea helps traders interpret model outputs and recognise their limits.

## The idea

Suppose you have data and a model with unknown parameters. The likelihood function measures how probable the data would be for each possible parameter value.

```
L(θ) = P(data | θ) = Π f(x_i | θ)
log L(θ) = Σ log f(x_i | θ)
```

MLE chooses the θ that maximises L(θ). In practice, we maximise the log likelihood, which turns products into sums and is easier to work with.

## A simple example: a coin or a win rate

**Example: Estimating a win rate by maximum likelihood**
A strategy wins 33 of 60 trades. Model each trade as a win with probability p. The likelihood is:

L(p) = p^33 × (1 minus p)^27.

The value of p that maximises this is p = 33 / 60 = 0.55, exactly the observed win rate. For this simple case, MLE gives the obvious answer. Its power comes in complex models where the best estimates are not obvious.

## Normal distribution

For data assumed to be normal, MLE gives:

```
μ̂ = sample mean
σ̂² = Σ (x - μ̂)² / n
```

Note the n rather than n minus 1: the MLE of variance is slightly biased in small samples, which is why the sample variance formula uses n minus 1. See [Variance and Standard Deviation](https://learn.tradelabsai.com/math/variance-and-standard-deviation/).

## MLE in finance

| Application | What is estimated | Lesson |
|---|---|---|
| GARCH volatility models | Parameters controlling how volatility reacts and persists | [GARCH](https://learn.tradelabsai.com/math/garch/) |
| Fat tailed distributions | Degrees of freedom of a Student t distribution | [Student's t-Distribution](https://learn.tradelabsai.com/math/students-t-distribution/) |
| ARIMA models | Time series coefficients | [ARIMA](https://learn.tradelabsai.com/math/arima/) |
| Logistic regression | Probability of an event, such as default or a price rise | [Regression and Classification Models](https://learn.tradelabsai.com/machine-learning/classification-models/) |
| Regime switching models | Probabilities and parameters of hidden regimes | [Structural Breaks and Regime Changes](https://learn.tradelabsai.com/math/regime-changes/) |
| Option model calibration | Parameters of stochastic volatility models (often with related methods) | [Stochastic Volatility and the Heston Model](https://learn.tradelabsai.com/options/heston-model/) |

## Why MLE is popular

- **General:** works for almost any model with a defined probability distribution.
- **Efficient:** in large samples, MLE estimates have the smallest possible variance under standard conditions.
- **Standard errors:** the curvature of the log likelihood gives estimates of uncertainty.
- **Model comparison:** likelihood based criteria such as AIC and BIC compare models while penalising complexity.

```
AIC = 2k - 2 log L
BIC = k log(n) - 2 log L
```

where k is the number of parameters and n the number of observations. Lower values are better.

## Limits and pitfalls

| Issue | Explanation |
|---|---|
| Wrong model | MLE finds the best parameters for the chosen model, even if the model is wrong |
| Small samples | Estimates can be biased and unstable |
| Local maxima | Complex likelihoods can have several peaks; optimisers may get stuck. See [Optimization](https://learn.tradelabsai.com/math/optimization/) |
| Overfitting | More parameters always increase the likelihood; use AIC, BIC or out of sample tests |
| Fat tails and outliers | Normal based MLE is sensitive to extreme values |
| Changing markets | Parameters estimated on the past may not hold. See [Structural Breaks and Regime Changes](https://learn.tradelabsai.com/math/regime-changes/) |

## MLE vs Bayesian estimation

MLE uses only the data. Bayesian estimation combines the likelihood with a prior distribution, which can stabilise estimates when data is limited. With lots of data, the two usually agree. See [Bayesian Statistics](https://learn.tradelabsai.com/math/bayesian-statistics/).

## Frequently asked questions

### What is maximum likelihood estimation?

A method that chooses model parameters that make the observed data most probable under the model.

### Where is maximum likelihood used in trading?

In estimating volatility models like GARCH, fitting return distributions, time series models, logistic regressions and calibrating some pricing models.

### What are the limits of maximum likelihood?

It assumes the model is correct, can be unstable in small samples, may find local rather than global maxima and can overfit with many parameters.

Next, learn the Bayesian approach in [Bayesian Statistics](https://learn.tradelabsai.com/math/bayesian-statistics/).

## Continue learning

- Next lesson: [Bayesian Statistics](https://learn.tradelabsai.com/math/bayesian-statistics/)
- Previous lesson: [Regression Analysis](https://learn.tradelabsai.com/math/regression-analysis/)
- Related: [Regression Analysis](https://learn.tradelabsai.com/math/regression-analysis/): Regression models how one variable relates to others. Learn linear regression, beta, R squared, multiple regression for factors, hedge ratios and common pitfalls.
- Related: [GARCH](https://learn.tradelabsai.com/math/garch/): GARCH models capture volatility clustering, where big moves follow big moves. Learn the GARCH(1,1) formula, persistence, forecasting and uses in risk and options.
- Related: [Probability Distributions Explained](https://learn.tradelabsai.com/math/probability-distributions/): Probability distributions describe the range and likelihood of outcomes. Learn the main ones used in trading, their shapes and when each applies.
- Related: [Bayesian Statistics](https://learn.tradelabsai.com/math/bayesian-statistics/): Bayesian statistics combines prior beliefs with data to estimate uncertain quantities. Learn priors and posteriors, shrinkage, credible intervals and trading uses.
- Related: [Optimization](https://learn.tradelabsai.com/math/optimization/): Optimisation finds inputs that maximise or minimise an objective, from portfolio weights to strategy settings. Learn the methods and how to avoid overfitting.
