TradeLabs AILearn

Regression Analysis

Regression models how one variable relates to others. Learn linear regression, beta, R squared, multiple regression for factors, hedge ratios and common pitfalls.

Intermediate3 min readUpdated 3 Oct 2026
Markdown
Lesson 19 of 46

Regression analysis estimates how one variable, such as a stock's return, relates to one or more other variables, such as the market's return or a set of factors. It is one of the most widely used tools in finance: it produces betas, hedge ratios, factor exposures and alpha estimates, and it underpins many forecasting models. Understanding how regression works, and how it can mislead, is essential for quantitative trading and risk management.

Simple linear regression#

y = α + β × x + ε
TermMeaning
yDependent variable (for example, a stock's return)
xIndependent variable (for example, the market's return)
α (intercept)Value of y when x is zero; in finance, often called alpha
β (slope)Change in y for a one unit change in x
ε (error)The part of y not explained by x

Ordinary least squares (OLS) chooses α and β to minimise the sum of squared errors.

β = Cov(x, y) / Var(x)

Worked example: market beta#

Multiple regression#

With several explanatory variables:

y = α + β1 x1 + β2 x2 + ... + βk xk + ε

Factor models use multiple regression to measure exposure to market, size, value, momentum and other factors. A fund's "alpha" is often defined as the intercept after controlling for these factors. See Factor Models.

Uses in trading#

UseExampleLesson
Beta and hedgingHedge a stock with index futuresAlpha and Beta
Hedge ratios for pairsRegress one stock's price on another'sPairs Trading
Factor exposuresMeasure a portfolio's tilt to value or momentumFactor Models
Performance attributionSeparate skill from factor returnsP&L and Performance Attribution
ForecastingPredict returns from signalsSignal Discovery
Cost modelsEstimate market impact from order sizeMarket Impact

Assumptions and checks#

AssumptionProblem if violatedRemedy
Linear relationshipBiased estimatesTransform variables, add terms
Independent errorsUnderstated standard errorsNewey West standard errors. See Autocorrelation and Partial Autocorrelation
Constant error varianceUnreliable inferenceRobust (heteroskedasticity consistent) standard errors
No strong multicollinearityUnstable coefficientsRemove or combine correlated variables
Stationary dataSpurious resultsUse returns, not prices; test for cointegration. See Stationarity, Differencing and Unit Roots
No extreme outliersDistorted fitRobust regression. See Outliers and Robust Statistics

Spurious regression#

Regressing one trending price series on another can produce high R² and "significant" coefficients even when the series are unrelated. Clive Granger and Paul Newbold showed this in 1974. Use returns or test for cointegration when working with prices. See Cointegration.

Overfitting#

Adding more variables always increases in sample R², even if they are noise. Use adjusted R², information criteria, regularisation (such as ridge and lasso regression) and out of sample tests to avoid overfitting. See Overfitting and Curve Fitting and Model Evaluation and Cross-Validation.

Frequently asked questions#

What is regression analysis in trading?#

A statistical method that estimates how one variable, such as a stock's return, relates to others, such as market or factor returns.

What does beta mean in a regression?#

The slope coefficient: how much the dependent variable changes, on average, for a one unit change in the independent variable.

What is a spurious regression?#

A regression that shows a strong relationship between unrelated variables, often caused by regressing trending, non stationary series on each other.

Next, learn a general method for fitting models in Maximum Likelihood.

Check your understanding

3 quick questions on this lesson. Get them all right to finish it.

Turn on JavaScript to take the quiz.

Finished this lesson?Sign in to save your progress across devices.
Next lessonMaximum LikelihoodMaximum likelihood estimation finds the model parameters that make observed data most probable. Learn the idea, simple examples, its use in GARCH and its limits.

Mentioned in