# Outliers and Robust Statistics

> Outliers can distort averages, correlations and backtests. Learn how to detect them, robust measures like the median and MAD, winsorising and when they matter.

Source: https://learn.tradelabsai.com/math/outliers-and-robust-statistics/  
Track: Math and Statistics · Level: Intermediate · Updated: 2026-10-03  
Publisher: TradeLabs AI (https://tradelabsai.com). Education, not financial advice.  
Cite as: TradeLabs Learn, "Outliers and Robust Statistics", https://learn.tradelabsai.com/math/outliers-and-robust-statistics/

An outlier is a value far from the rest of the data. In trading, outliers come in two kinds: errors, such as a bad tick that prints a price 50% away from the market, and genuine extreme events, such as a crash day or a takeover jump. Errors should be fixed or removed; genuine extremes often matter most for risk and must be kept. Robust statistics are methods that are less sensitive to outliers, helping researchers see the typical pattern without letting a few values dominate.

## How outliers distort common statistics

| Statistic | Sensitivity to outliers |
|---|---|
| Mean | High: one extreme value shifts it |
| Standard deviation | Very high: squared deviations amplify extremes |
| Pearson correlation | High: a few points can create or hide a relationship |
| Ordinary regression | High: outliers pull the fitted line. See [Regression Analysis](https://learn.tradelabsai.com/math/regression-analysis/) |
| Median | Low |
| Median absolute deviation (MAD) | Low |
| Spearman rank correlation | Low |

**Example: One bad tick**
A stock's daily closes hover around $100, with a typical daily move of 1%. A data error records one close at $10, then $100 the next day. That single error creates returns of minus 90% and +900%, inflating the standard deviation enormously and wrecking any volatility based indicator or backtest. A median based volatility estimate would barely change. See [Cleaning Market Data](https://learn.tradelabsai.com/programming/cleaning-market-data/).

## Detecting outliers

| Method | Rule of thumb |
|---|---|
| Z score | Values more than 3 to 4 standard deviations from the mean. See [Percentiles, Quantiles and Z-Scores](https://learn.tradelabsai.com/math/z-scores/) |
| Robust z score | (x minus median) / (1.4826 × MAD), flag values beyond about 3.5 |
| Interquartile range (IQR) | Below Q1 minus 1.5 × IQR or above Q3 + 1.5 × IQR |
| Domain checks | Prices outside the day's trading range, negative volumes, impossible jumps |
| Visual inspection | Charts and histograms |

The standard z score method is itself distorted by outliers, since the mean and standard deviation include them. Robust z scores avoid this.

## Robust measures

```
MAD = median(|x - median(x)|)
robust standard deviation ≈ 1.4826 × MAD   (for normal data)
```

Other robust tools:

- **Trimmed mean:** drop a percentage of the highest and lowest values, then average.
- **Winsorising:** cap extreme values at chosen percentiles, such as the 1st and 99th.
- **Rank based methods:** Spearman correlation, rank transformed signals.
- **Robust regression:** methods such as Huber regression that reduce outliers' influence.

## Winsorising signals in quant research

Quantitative strategies often winsorise or rank signals across stocks before combining them. For example, a company with a tiny positive earnings figure might show a P/E of 5,000, which would dominate a raw value signal. Capping extremes or using ranks keeps a few companies from distorting the portfolio. See [Combining Signals](https://learn.tradelabsai.com/research/combining-signals/) and [Factor Investing Explained](https://learn.tradelabsai.com/research/factor-investing-explained/).

## When outliers are the point

In risk management, extreme values are not noise to be removed; they are the main concern. Tail events drive drawdowns, margin calls and ruin. Removing genuine crash days from a backtest makes a strategy look far safer than it is. See [Fat Tails](https://learn.tradelabsai.com/math/fat-tails/) and [Maximum Drawdown](https://learn.tradelabsai.com/portfolio/maximum-drawdown/).

**Example: Results driven by a few days**
Studies of stock market returns have found that missing a small number of the best days dramatically reduces long term returns, and that many of the best days occur close to the worst days, during volatile periods. A strategy's performance can likewise depend on a handful of trades. Check how results change if the top and bottom 1% of trades are removed; if the edge disappears, it depends on outliers. See [Robustness and Stress Testing](https://learn.tradelabsai.com/research/robustness-and-stress-testing/).

## Practical guidelines

1. **Separate errors from genuine events** using domain knowledge and multiple data sources.
2. **Fix or remove errors;** keep genuine extremes for risk analysis.
3. **Use robust measures** for signals and typical behaviour.
4. **Report sensitivity:** show results with and without extreme values.
5. **Document every adjustment** to keep research reproducible. See [Backtest Reproducibility](https://learn.tradelabsai.com/research/backtest-reproducibility/).

## Frequently asked questions

### What is an outlier in trading data?

A value far from the rest of the data, caused either by a data error or by a genuine extreme market event.

### What are robust statistics?

Statistical methods, such as the median and median absolute deviation, that are less affected by outliers than the mean and standard deviation.

### Should I remove outliers from a backtest?

Remove data errors, but keep genuine extreme events, since they are often the most important for understanding risk.

Next, learn what makes a result convincing in [Statistical Significance in Trading](https://learn.tradelabsai.com/math/statistical-significance/).

## Continue learning

- Next lesson: [Statistical Significance in Trading](https://learn.tradelabsai.com/math/statistical-significance/)
- Previous lesson: [Bootstrap and Permutation Tests](https://learn.tradelabsai.com/math/bootstrap-and-permutation-tests/)
- Related: [Bootstrap and Permutation Tests](https://learn.tradelabsai.com/math/bootstrap-and-permutation-tests/): Bootstrap and permutation tests use resampling to measure uncertainty and test significance without strict assumptions. Learn how they work, examples and pitfalls.
- Related: [Mean, Median and Mode](https://learn.tradelabsai.com/math/mean-median-and-mode/): The mean, median and mode measure the centre of data in different ways. Learn when each is best for trading data, how outliers distort averages and geometric means.
- Related: [Fat Tails](https://learn.tradelabsai.com/math/fat-tails/): Fat tails mean extreme market moves happen far more often than the normal curve predicts. Learn the evidence, the causes, how to measure them and how to manage them.
- Related: [Cleaning Market Data](https://learn.tradelabsai.com/programming/cleaning-market-data/): Raw market data contains bad ticks, gaps, duplicates and wrong timestamps. Learn how to detect and fix common data errors without distorting your backtests.
- Related: [Percentiles, Quantiles and Z-Scores](https://learn.tradelabsai.com/math/z-scores/): A z score shows how many standard deviations a value is from its mean. Learn the formula, its uses in mean reversion and pairs trading, and the pitfalls.
- Related: [Regression Analysis](https://learn.tradelabsai.com/math/regression-analysis/): Regression models how one variable relates to others. Learn linear regression, beta, R squared, multiple regression for factors, hedge ratios and common pitfalls.
