Supervised vs Unsupervised Learning
Supervised learning trains models on examples with known answers. Learn regression versus classification, how to define trading targets and labels, and key pitfalls.
Supervised learning is the most widely used form of machine learning. You give a model many examples, each with inputs (features) and a known answer (the label or target), and it learns a function that maps inputs to answers. In trading, the inputs might be recent returns, volatility, valuation ratios or sentiment scores, and the target might be next week's return or whether a trade hit its profit target first. The model is only as useful as the question you ask it, which makes defining the target one of the most important decisions.
Regression versus classification#
| Regression | Classification | |
|---|---|---|
| Predicts | A number | A category |
| Trading examples | Next month's return, tomorrow's volatility | Up or down, hit target or stop first, regime |
| Typical models | Linear regression, ridge, gradient boosting | Logistic regression, tree ensembles, neural networks |
| Evaluation | Error measures, correlation with outcomes | Accuracy, precision, recall, log loss |
| Lesson | Regression Analysis | Regression and Classification Models |
Defining the target#
| Target | Pros | Cons |
|---|---|---|
| Next period return | Direct link to profit | Very noisy |
| Return sign (up or down) | Simple | Ignores size; tiny moves count the same as big ones |
| Return above a threshold | Focuses on meaningful moves | Fewer positive examples |
| Triple barrier label | Matches how trades exit: target, stop or time limit | More complex to compute |
| Volatility | More predictable | Not directly a direction trade |
| Cross sectional rank | Predicts which assets beat others | Needs many assets |
The triple barrier method, popularised by Marcos López de Prado, labels each event by which comes first: an upper barrier (profit target), a lower barrier (stop loss) or a vertical barrier (time limit). It ties labels to realistic trade outcomes. See Profit Targets and Stop Loss Strategies.
The supervised workflow#
- Collect point in time features and targets. See Point-in-Time and Survivorship-Free Data.
- Align carefully: features must be known before the target period starts.
- Split data by time into training, validation and test sets.
- Train simple models first, such as linear or logistic regression.
- Tune on validation data, never on the test set.
- Evaluate once on the test set, then translate predictions into trades with costs.
Overlapping labels#
If each label covers the next 10 days and you create a label every day, consecutive labels overlap heavily and share most of their information. This makes the data look larger than it is and leaks information across training and test splits. Remedies include sampling less often, weighting samples by uniqueness and purging overlapping samples near split boundaries. See Model Evaluation and Cross-Validation.
Class imbalance#
When positive examples are rare, such as large moves, a model can score high accuracy by always predicting the common class. Use balanced metrics, class weights or resampling, and judge the model by trading results. See Regression and Classification Models.
From prediction to position#
A prediction is not a trade. Decide how predictions become positions: trade only when confidence is high, size positions by predicted strength, or rank assets and go long the top and short the bottom. Each choice changes turnover and costs. See Position Sizing and Signal Turnover, Breadth and Neutralization.
Frequently asked questions#
What is supervised learning in trading?#
Training a model on historical examples with known outcomes, such as features and the following return, so it can predict outcomes for new data.
What is the triple barrier method?#
A labelling approach that marks each event by whether price first hits a profit target, a stop loss or a time limit, matching how real trades exit.
Should I predict returns or direction?#
Both are used. Direction is simpler but ignores move size; returns or volatility scaled labels often relate more closely to profit.
Next, learn how classification models make decisions in Regression and Classification Models.
3 quick questions on this lesson. Get them all right to finish it.
Turn on JavaScript to take the quiz.
Mentioned in
- Machine Learning in TradingMachine Learning