Feature Engineering
Features are the inputs that give trading models a chance. Learn the main feature families, how to make them stationary and comparable, and how to avoid leakage.
In trading machine learning, the choice of inputs usually matters more than the choice of model. Feature engineering means turning raw data, such as prices, volumes, financial statements or news, into variables that capture something meaningful and that a model can learn from. Good features are grounded in a reason the market might behave a certain way, are calculated using only past information and are scaled so they mean the same thing across time and across assets. A simple model with thoughtful features almost always beats a complex model fed raw prices.
Main feature families#
| Family | Examples | Lesson |
|---|---|---|
| Momentum and trend | Past returns over 1, 3 and 12 months; distance from moving averages | Momentum Factor |
| Mean reversion | Short term return, z score versus a rolling mean, RSI | Mean Reversion |
| Volatility | Realised volatility, ATR, implied volatility, volatility ratios | Historical and Realized Volatility |
| Volume and liquidity | Relative volume, turnover, spread, order imbalance | Relative Volume |
| Fundamentals | Valuation ratios, profitability, growth, revisions | Value Factor |
| Calendar | Day of week, month, time to earnings or expiry | Seasonality in Commodities |
| Cross asset | Rates, dollar, commodities, sector returns | Currency Correlations |
| Sentiment and alternative | News tone, social activity, web traffic | Sentiment Data |
Make features stationary#
Raw prices trend and change scale over time: a stock at $20 in 2015 and $200 today. A model that learns from price levels learns nothing transferable. Convert to forms with stable statistical properties:
- Returns instead of prices.
- Ratios such as price divided by its 50 day average.
- Changes in rates or spreads.
- Volatility scaled values, such as a return divided by recent volatility.
See Stationarity, Differencing and Unit Roots.
Make features comparable#
| Technique | Purpose |
|---|---|
| Rolling z score | Express a value relative to its own recent history. See Percentiles, Quantiles and Z-Scores |
| Cross sectional rank | Rank each asset against others on the same date, from 0 to 1 |
| Volatility scaling | Put calm and volatile assets on one scale |
| Winsorising | Cap extreme values to limit outlier influence. See Outliers and Robust Statistics |
Use only past data for any scaling. A z score computed with a full sample mean leaks the future. See Rolling and Expanding Windows.
Avoiding leakage#
| Leak | Example | Fix |
|---|---|---|
| Same bar data | Using today's close to predict today's return | Shift features by one period |
| Full sample statistics | Normalising with mean and standard deviation of all data | Use rolling or expanding windows |
| Restated fundamentals | Using revised earnings | Point in time data. See Point-in-Time and Survivorship-Free Data |
| Survivorship | Features only for stocks that still exist | Survivorship free universe |
| Target in disguise | A feature computed over the target period | Check every feature's time window |
See Data Leakage and Look-Ahead Bias.
How many features?#
More features are not always better. Each extra feature adds a chance of fitting noise, and many features are highly correlated. Start with a few well motivated ones, check their individual predictive power out of sample, and add more only if they improve validation results. Techniques such as principal component analysis or clustering can group correlated features. See Overfitting and Curve Fitting.
Testing a single feature#
Before modelling, test each feature alone: sort assets or periods into groups by the feature and compare later returns, or compute the information coefficient, the rank correlation between the feature and future returns. A feature that shows no relationship alone rarely becomes useful inside a model. See Signal Discovery.
Frequently asked questions#
What is feature engineering in trading?#
Creating model inputs from raw data, such as returns, volatility, valuation ratios or sentiment scores, in forms a model can learn from.
Why should features be stationary?#
Non stationary inputs such as raw prices change scale over time, so patterns learned in one period do not apply in another.
How do I avoid look ahead bias in features?#
Use only data available before each prediction, shift features appropriately, use rolling statistics and use point in time fundamental data.
Next, learn how to validate models properly in Model Evaluation and Cross-Validation.
3 quick questions on this lesson. Get them all right to finish it.
Turn on JavaScript to take the quiz.
Mentioned in
- Machine Learning in TradingMachine Learning
- Random Forests and Gradient BoostingMachine Learning
- Neural Networks and Deep LearningMachine Learning
- Granger CausalityMath and Statistics
- Signal DiscoveryResearch and Backtesting