TradeLabs AILearn

Feature Engineering

Features are the inputs that give trading models a chance. Learn the main feature families, how to make them stationary and comparable, and how to avoid leakage.

Advanced3 min readUpdated 3 Oct 2026
Markdown
Lesson 6 of 10

In trading machine learning, the choice of inputs usually matters more than the choice of model. Feature engineering means turning raw data, such as prices, volumes, financial statements or news, into variables that capture something meaningful and that a model can learn from. Good features are grounded in a reason the market might behave a certain way, are calculated using only past information and are scaled so they mean the same thing across time and across assets. A simple model with thoughtful features almost always beats a complex model fed raw prices.

Main feature families#

FamilyExamplesLesson
Momentum and trendPast returns over 1, 3 and 12 months; distance from moving averagesMomentum Factor
Mean reversionShort term return, z score versus a rolling mean, RSIMean Reversion
VolatilityRealised volatility, ATR, implied volatility, volatility ratiosHistorical and Realized Volatility
Volume and liquidityRelative volume, turnover, spread, order imbalanceRelative Volume
FundamentalsValuation ratios, profitability, growth, revisionsValue Factor
CalendarDay of week, month, time to earnings or expirySeasonality in Commodities
Cross assetRates, dollar, commodities, sector returnsCurrency Correlations
Sentiment and alternativeNews tone, social activity, web trafficSentiment Data

Make features stationary#

Raw prices trend and change scale over time: a stock at $20 in 2015 and $200 today. A model that learns from price levels learns nothing transferable. Convert to forms with stable statistical properties:

  • Returns instead of prices.
  • Ratios such as price divided by its 50 day average.
  • Changes in rates or spreads.
  • Volatility scaled values, such as a return divided by recent volatility.

See Stationarity, Differencing and Unit Roots.

Make features comparable#

TechniquePurpose
Rolling z scoreExpress a value relative to its own recent history. See Percentiles, Quantiles and Z-Scores
Cross sectional rankRank each asset against others on the same date, from 0 to 1
Volatility scalingPut calm and volatile assets on one scale
WinsorisingCap extreme values to limit outlier influence. See Outliers and Robust Statistics

Use only past data for any scaling. A z score computed with a full sample mean leaks the future. See Rolling and Expanding Windows.

Avoiding leakage#

LeakExampleFix
Same bar dataUsing today's close to predict today's returnShift features by one period
Full sample statisticsNormalising with mean and standard deviation of all dataUse rolling or expanding windows
Restated fundamentalsUsing revised earningsPoint in time data. See Point-in-Time and Survivorship-Free Data
SurvivorshipFeatures only for stocks that still existSurvivorship free universe
Target in disguiseA feature computed over the target periodCheck every feature's time window

See Data Leakage and Look-Ahead Bias.

How many features?#

More features are not always better. Each extra feature adds a chance of fitting noise, and many features are highly correlated. Start with a few well motivated ones, check their individual predictive power out of sample, and add more only if they improve validation results. Techniques such as principal component analysis or clustering can group correlated features. See Overfitting and Curve Fitting.

Testing a single feature#

Before modelling, test each feature alone: sort assets or periods into groups by the feature and compare later returns, or compute the information coefficient, the rank correlation between the feature and future returns. A feature that shows no relationship alone rarely becomes useful inside a model. See Signal Discovery.

Frequently asked questions#

What is feature engineering in trading?#

Creating model inputs from raw data, such as returns, volatility, valuation ratios or sentiment scores, in forms a model can learn from.

Why should features be stationary?#

Non stationary inputs such as raw prices change scale over time, so patterns learned in one period do not apply in another.

How do I avoid look ahead bias in features?#

Use only data available before each prediction, shift features appropriately, use rolling statistics and use point in time fundamental data.

Next, learn how to validate models properly in Model Evaluation and Cross-Validation.

Check your understanding

3 quick questions on this lesson. Get them all right to finish it.

Turn on JavaScript to take the quiz.

Finished this lesson?Sign in to save your progress across devices.
Next lessonModel Evaluation and Cross-ValidationStandard cross validation leaks future data in time series. Learn time series splits, purging and embargo, combinatorial purged cross validation and good practice.

Mentioned in