TradeLabs AILearn

Random Forests and Gradient Boosting

Random forests and gradient boosted trees are strong models for tabular trading data. Learn how they work, key settings, feature importance and overfitting risks.

Advanced3 min readUpdated 3 Oct 2026
Markdown
Lesson 4 of 10

Tree ensemble models are among the most effective machine learning tools for the kind of data traders typically use: tables of features such as returns, volatility, valuation ratios and sentiment scores. A random forest builds many decision trees on random variations of the data and averages them. Gradient boosting builds trees one after another, each correcting the errors of those before it. Both capture non linear relationships and interactions between features without heavy preparation, which makes them popular starting points for quantitative research.

From one tree to a forest#

A single decision tree splits data with simple questions: is the 20 day return above 2%? Is volatility below 15%? It is easy to understand but tends to memorise training data. A random forest reduces this by:

  1. Bootstrapping: each tree trains on a random sample of rows drawn with replacement.
  2. Feature subsampling: each split considers only a random subset of features.
  3. Averaging: predictions from hundreds of trees are averaged or voted.

The randomness makes trees different from each other, and averaging different errors cancels much of the noise.

Gradient boosting#

Random forestGradient boosting
Tree buildingIndependent, in parallelSequential, each fixes previous errors
Typical tree depthDeepShallow
Overfitting riskLower with defaultsHigher; needs careful tuning
Accuracy on tabular dataStrongOften the strongest
Popular librariesscikit-learnXGBoost, LightGBM, CatBoost

Key settings#

SettingEffect
Number of treesMore trees reduce variance, with diminishing returns
Maximum depthDeeper trees capture more complexity and overfit more
Minimum samples per leafHigher values smooth predictions; useful for noisy financial data
Features per splitFewer features add randomness
Learning rate (boosting)Smaller rates need more trees but generalise better
Early stopping (boosting)Stops adding trees when validation error stops improving

For noisy market data, shallow trees and large minimum leaf sizes usually work better than the defaults designed for cleaner problems.

Feature importance#

Tree models report which features they relied on. Two common methods:

  • Impurity based importance: how much each feature reduced error in splits. Fast, but biased toward features with many unique values.
  • Permutation importance: how much performance drops when a feature's values are shuffled. More reliable, especially when measured on out of sample data.

Importance shows what the model used, not that the relationship is real or causal. See Feature Engineering.

Why trees suit financial features#

  • No need to scale features.
  • Robust to outliers in inputs, since splits depend on order, not magnitude. See Outliers and Robust Statistics.
  • Capture interactions, such as momentum working only in low volatility regimes.
  • Handle mixed feature types.

Limitations#

  • Cannot extrapolate: predictions stay within the range of training targets.
  • Can still overfit noisy data, especially boosting with deep trees.
  • Struggle with raw sequences compared with specialised models; engineered features work better.
  • Non stationarity means relationships learned years ago may not hold. See Structural Breaks and Regime Changes.

Validation#

Always validate with time aware methods, such as walk forward or purged cross validation, never random shuffles. See Model Evaluation and Cross-Validation and Walk-Forward Validation and Preventing Overfitting.

Frequently asked questions#

Are random forests good for trading?#

They are a solid, robust choice for tabular features, though their predictions must still be validated carefully and translated into trades with costs.

What is the difference between random forests and gradient boosting?#

Random forests average many independent deep trees; gradient boosting builds shallow trees sequentially, each correcting earlier errors.

What does feature importance tell me?#

Which features the model relied on most. It does not prove a real or causal relationship and can reveal data leakage.

Next, learn how neural networks work and where they fit in Neural Networks and Deep Learning.

Check your understanding

3 quick questions on this lesson. Get them all right to finish it.

Turn on JavaScript to take the quiz.

Finished this lesson?Sign in to save your progress across devices.
Next lessonNeural Networks and Deep LearningNeural networks power deep learning, from LSTMs to transformers. Learn how they work, where they help in trading, especially with text and images, and their risks.

Mentioned in