# Regression and Classification Models

> How classification models predict up or down moves and trade outcomes. Learn logistic regression, probability thresholds, precision, recall and confusion matrices.

Source: https://learn.tradelabsai.com/machine-learning/classification-models/  
Track: Machine Learning · Level: Advanced · Updated: 2026-10-03  
Publisher: TradeLabs AI (https://tradelabsai.com). Education, not financial advice.  
Cite as: TradeLabs Learn, "Regression and Classification Models", https://learn.tradelabsai.com/machine-learning/classification-models/

Classification models answer questions with categories: will this stock go up or down tomorrow, will this trade reach its target before its stop, is the market in a calm or volatile regime? Most classifiers output a probability for each class, which makes them natural tools for trading decisions: trade only when the model is confident enough. Getting value from a classifier requires choosing the right evaluation measures, because accuracy alone often misleads.

## Common classification models

| Model | How it works | Strengths |
|---|---|---|
| Logistic regression | Weighted sum of features passed through an S shaped curve to give a probability | Simple, interpretable, hard to overfit |
| Decision tree | Series of yes or no splits on features | Captures interactions; easy to explain; overfits alone |
| Random forest | Many trees on random subsets, averaged | Robust, handles non linear patterns. See [Random Forests and Gradient Boosting](https://learn.tradelabsai.com/machine-learning/random-forests/) |
| Gradient boosted trees | Trees built one after another to fix previous errors | Often the strongest for tabular data |
| Neural networks | Layers of weighted connections | Flexible; need more data. See [Neural Networks and Deep Learning](https://learn.tradelabsai.com/machine-learning/neural-networks/) |
| Support vector machines | Find a boundary that separates classes with the widest margin | Effective in some medium sized problems |

## The confusion matrix

| | Predicted up | Predicted down |
|---|---|---|
| Actually up | True positive | False negative |
| Actually down | False positive | True negative |

From this table come the key measures:

- **Accuracy:** share of all predictions that were correct.
- **Precision:** of the times the model said up, how often it was right.
- **Recall:** of all the actual up moves, how many the model caught.
- **Log loss:** how good the predicted probabilities are, not just the labels.

See [Statistical Power and Type I and II Errors](https://learn.tradelabsai.com/math/type-i-and-ii-errors/).

**Example: Why precision can matter more than accuracy**
A model predicts whether a stock will rise more than 3% within five days. Out of 1,000 test cases, 100 actually rise that much. The model flags 50 cases as positive, and 30 of them do rise more than 3%. Precision is 30 out of 50, or 60%; recall is 30 out of 100, or 30%; accuracy is (30 correct positives plus 880 correct negatives) out of 1,000, or 91%. The 91% accuracy is mostly from correctly saying "no" to common cases. For trading, the 60% precision on the trades actually taken, combined with the win and loss sizes, determines profit. See [Expectancy](https://learn.tradelabsai.com/risk/expectancy/).

## Choosing a probability threshold

Classifiers typically output a probability, and the default rule is to predict "up" above 50%. In trading you can raise the threshold, trading only when the model is more confident. Higher thresholds usually mean fewer trades with higher precision, up to a point. Choose the threshold on validation data, based on expected profit after costs rather than accuracy.

| Threshold | Trades | Precision (illustrative) |
|---|---|---|
| 0.50 | Many | 52% |
| 0.55 | Fewer | 55% |
| 0.60 | Few | 58% |

## Calibration

A well calibrated model's 60% predictions come true about 60% of the time. Many models, especially tree ensembles, produce poorly calibrated probabilities. Calibration methods such as Platt scaling or isotonic regression fix this, which matters if you size positions by probability. See [Kelly Criterion](https://learn.tradelabsai.com/risk/kelly-criterion/).

## Pitfalls

1. **Class imbalance:** rare events make accuracy meaningless; use precision, recall and balanced weights.
2. **Leakage:** features calculated with future data. See [Data Leakage](https://learn.tradelabsai.com/research/data-leakage/).
3. **Random train and test splits** on time series, which leak future information. See [Model Evaluation and Cross-Validation](https://learn.tradelabsai.com/machine-learning/cross-validation/).
4. **Ignoring move size:** a model right on small moves and wrong on large ones can lose money with high accuracy.
5. **Too many features** relative to data. See [Overfitting and Curve Fitting](https://learn.tradelabsai.com/research/overfitting-and-curve-fitting/).

## Frequently asked questions

### What is a classification model in trading?

A model that predicts a category, such as up or down or target hit versus stop hit, usually with a probability for each outcome.

### Is accuracy a good measure for trading models?

Not on its own. Precision on trades taken, the size of wins and losses, and costs determine profitability.

### What is model calibration?

How well predicted probabilities match actual frequencies; a calibrated model's 70% predictions come true about 70% of the time.

Next, learn one of the most reliable model types in [Random Forests and Gradient Boosting](https://learn.tradelabsai.com/machine-learning/random-forests/).

## Continue learning

- Next lesson: [Random Forests and Gradient Boosting](https://learn.tradelabsai.com/machine-learning/random-forests/)
- Previous lesson: [Supervised vs Unsupervised Learning](https://learn.tradelabsai.com/machine-learning/supervised-learning/)
- Related: [Supervised vs Unsupervised Learning](https://learn.tradelabsai.com/machine-learning/supervised-learning/): Supervised learning trains models on examples with known answers. Learn regression versus classification, how to define trading targets and labels, and key pitfalls.
- Related: [Random Forests and Gradient Boosting](https://learn.tradelabsai.com/machine-learning/random-forests/): Random forests and gradient boosted trees are strong models for tabular trading data. Learn how they work, key settings, feature importance and overfitting risks.
- Related: [Neural Networks and Deep Learning](https://learn.tradelabsai.com/machine-learning/neural-networks/): Neural networks power deep learning, from LSTMs to transformers. Learn how they work, where they help in trading, especially with text and images, and their risks.
- Related: [Statistical Power and Type I and II Errors](https://learn.tradelabsai.com/math/type-i-and-ii-errors/): Type I errors are false positives; type II errors are missed real effects. Learn how they apply to strategy testing, the trade off between them and power.
- Related: [Expectancy](https://learn.tradelabsai.com/risk/expectancy/): Expectancy is the average amount you win or lose per trade. Learn the formula, how win rate and payoff combine, expectancy in R and how to improve it.
