Regression and Classification Models
How classification models predict up or down moves and trade outcomes. Learn logistic regression, probability thresholds, precision, recall and confusion matrices.
Classification models answer questions with categories: will this stock go up or down tomorrow, will this trade reach its target before its stop, is the market in a calm or volatile regime? Most classifiers output a probability for each class, which makes them natural tools for trading decisions: trade only when the model is confident enough. Getting value from a classifier requires choosing the right evaluation measures, because accuracy alone often misleads.
Common classification models#
| Model | How it works | Strengths |
|---|---|---|
| Logistic regression | Weighted sum of features passed through an S shaped curve to give a probability | Simple, interpretable, hard to overfit |
| Decision tree | Series of yes or no splits on features | Captures interactions; easy to explain; overfits alone |
| Random forest | Many trees on random subsets, averaged | Robust, handles non linear patterns. See Random Forests and Gradient Boosting |
| Gradient boosted trees | Trees built one after another to fix previous errors | Often the strongest for tabular data |
| Neural networks | Layers of weighted connections | Flexible; need more data. See Neural Networks and Deep Learning |
| Support vector machines | Find a boundary that separates classes with the widest margin | Effective in some medium sized problems |
The confusion matrix#
| Predicted up | Predicted down | |
|---|---|---|
| Actually up | True positive | False negative |
| Actually down | False positive | True negative |
From this table come the key measures:
- Accuracy: share of all predictions that were correct.
- Precision: of the times the model said up, how often it was right.
- Recall: of all the actual up moves, how many the model caught.
- Log loss: how good the predicted probabilities are, not just the labels.
See Statistical Power and Type I and II Errors.
Choosing a probability threshold#
Classifiers typically output a probability, and the default rule is to predict "up" above 50%. In trading you can raise the threshold, trading only when the model is more confident. Higher thresholds usually mean fewer trades with higher precision, up to a point. Choose the threshold on validation data, based on expected profit after costs rather than accuracy.
| Threshold | Trades | Precision (illustrative) |
|---|---|---|
| 0.50 | Many | 52% |
| 0.55 | Fewer | 55% |
| 0.60 | Few | 58% |
Calibration#
A well calibrated model's 60% predictions come true about 60% of the time. Many models, especially tree ensembles, produce poorly calibrated probabilities. Calibration methods such as Platt scaling or isotonic regression fix this, which matters if you size positions by probability. See Kelly Criterion.
Pitfalls#
- Class imbalance: rare events make accuracy meaningless; use precision, recall and balanced weights.
- Leakage: features calculated with future data. See Data Leakage.
- Random train and test splits on time series, which leak future information. See Model Evaluation and Cross-Validation.
- Ignoring move size: a model right on small moves and wrong on large ones can lose money with high accuracy.
- Too many features relative to data. See Overfitting and Curve Fitting.
Frequently asked questions#
What is a classification model in trading?#
A model that predicts a category, such as up or down or target hit versus stop hit, usually with a probability for each outcome.
Is accuracy a good measure for trading models?#
Not on its own. Precision on trades taken, the size of wins and losses, and costs determine profitability.
What is model calibration?#
How well predicted probabilities match actual frequencies; a calibrated model's 70% predictions come true about 70% of the time.
Next, learn one of the most reliable model types in Random Forests and Gradient Boosting.
3 quick questions on this lesson. Get them all right to finish it.
Turn on JavaScript to take the quiz.
Mentioned in
- Maximum LikelihoodMath and Statistics