# Reinforcement Learning

> Reinforcement learning trains agents to act by rewarding good outcomes. Learn how it applies to trading and execution, how rewards are designed and why it is hard.

Source: https://learn.tradelabsai.com/machine-learning/reinforcement-learning/  
Track: Machine Learning · Level: Advanced · Updated: 2026-10-03  
Publisher: TradeLabs AI (https://tradelabsai.com). Education, not financial advice.  
Cite as: TradeLabs Learn, "Reinforcement Learning", https://learn.tradelabsai.com/machine-learning/reinforcement-learning/

Reinforcement learning (RL) is a branch of machine learning in which an agent learns to make decisions by acting in an environment and receiving rewards or penalties. It famously mastered games such as Go and chess. Trading looks like a natural fit: an agent observes the market, chooses to buy, sell or hold, and is rewarded with profit. In practice, RL in trading is difficult, because markets are noisy, change over time and cannot be replayed endlessly with your actions affecting them realistically. Its most promising uses are narrow, well defined problems such as order execution and market making.

## The RL framework

| Element | Meaning | Trading example |
|---|---|---|
| Agent | The decision maker | The trading algorithm |
| Environment | What the agent interacts with | The market or a simulator |
| State | What the agent observes | Prices, position, time remaining, order book |
| Action | What the agent can do | Buy, sell, hold, place a limit order at a price |
| Reward | Feedback after actions | Profit, minus costs and risk penalties |
| Policy | The agent's strategy | A mapping from states to actions |

## Main approaches

| Approach | Idea | Examples |
|---|---|---|
| Value based | Learn how good each action is in each state | Q learning, deep Q networks |
| Policy based | Learn the policy directly | Policy gradients, PPO |
| Actor critic | Combine both | A2C, SAC |
| Model based | Learn a model of the environment and plan | Used where simulators are reliable |

## Where RL fits best in trading

| Problem | Why RL suits it |
|---|---|
| Optimal execution | Clear goal (minimise cost), limited horizon, actions affect outcomes. See [Optimal Execution and the Almgren-Chriss Model](https://learn.tradelabsai.com/orders/optimal-execution/) |
| Market making | Repeated decisions on quote placement and inventory. See [Market Making](https://learn.tradelabsai.com/strategies/market-making/) |
| Hedging derivatives | Balancing hedging costs and risk, sometimes called deep hedging. See [Delta Hedging](https://learn.tradelabsai.com/options/delta-hedging/) |
| Portfolio rebalancing | Trading off tracking error and costs. See [Rebalancing](https://learn.tradelabsai.com/portfolio/rebalancing/) |

## Designing the reward

The reward shapes everything the agent learns. Rewarding raw profit encourages excessive risk. Better rewards include risk adjusted returns, penalties for large positions or drawdowns, and transaction costs. Poorly designed rewards produce agents that exploit flaws in the simulator rather than learning real skill.

**Example: An execution agent**
An agent must sell 10,000 shares within 30 minutes. Each minute it chooses how many shares to sell, observing the time remaining, shares left, the spread and recent volume. The reward is the negative of implementation shortfall: the difference between the arrival price and the average sale price, plus a penalty for any shares left at the end. Trained in a simulator calibrated to historical order book data, the agent learns to sell more when spreads are tight and volume is high, and to speed up as time runs out. Compared with a simple equal slices schedule, it might reduce cost modestly in simulation; the real test is live performance, since the simulator cannot fully model how other traders react. See [Implementation Shortfall](https://learn.tradelabsai.com/orders/implementation-shortfall/) and [VWAP, TWAP and POV Execution](https://learn.tradelabsai.com/orders/vwap-twap-and-pov-execution/).

## Why RL is hard in trading

- **Limited data:** you cannot replay the market millions of times like a game.
- **Simulators are imperfect:** historical replay ignores how your orders would have moved prices. See [Market Impact](https://learn.tradelabsai.com/orders/market-impact/).
- **Non stationarity:** the environment changes, so learned policies go stale. See [Structural Breaks and Regime Changes](https://learn.tradelabsai.com/math/regime-changes/).
- **Noisy rewards:** profit is dominated by randomness, making learning slow and unstable.
- **Overfitting:** agents easily memorise the training period. See [Overfitting and Curve Fitting](https://learn.tradelabsai.com/research/overfitting-and-curve-fitting/).
- **Safety:** exploration with real money is costly. See [Risk Controls and Kill Switches](https://learn.tradelabsai.com/algo-trading/risk-controls-and-kill-switches/).

## Practical advice

1. **Start with a narrow problem** with clear rewards, such as execution.
2. **Build a realistic simulator** and validate it against real outcomes.
3. **Compare with simple baselines,** such as TWAP or a rule based policy.
4. **Constrain actions** to safe ranges.
5. **Test out of sample** across different market conditions. See [Walk-Forward Validation and Preventing Overfitting](https://learn.tradelabsai.com/machine-learning/walk-forward-validation/).

## Frequently asked questions

### What is reinforcement learning in trading?

A machine learning method where an agent learns trading or execution decisions by acting in a market environment and receiving rewards such as risk adjusted profit.

### Does reinforcement learning work for trading?

It shows promise in narrow problems such as execution, market making and hedging; using it to predict and trade prices directly is much harder.

### Why is reward design important?

The agent optimises whatever reward it is given, so rewards must include costs and risk, or it will learn risky or simulator specific behaviour.

You have finished the Machine Learning track. Continue with how to measure performance in [Measuring Returns and CAGR](https://learn.tradelabsai.com/portfolio/measuring-returns-and-cagr/).

## Continue learning

- Previous lesson: [Online Learning](https://learn.tradelabsai.com/machine-learning/online-learning/)
- Related: [Online Learning](https://learn.tradelabsai.com/machine-learning/online-learning/): Online learning updates a model with each new observation instead of retraining in batches. Learn how it works, forgetting factors, drift detection and its risks.
- Related: [Machine Learning in Trading](https://learn.tradelabsai.com/machine-learning/machine-learning-in-trading/): An honest guide to machine learning in trading: where it helps, why it often fails on market data, the main model types and a sound workflow for using it safely.
- Related: [Optimal Execution and the Almgren-Chriss Model](https://learn.tradelabsai.com/orders/optimal-execution/): Optimal execution balances market impact against price risk when trading large orders. Learn the Almgren-Chriss model, its trade off and what it means in practice.
- Related: [Market Making](https://learn.tradelabsai.com/strategies/market-making/): Market making quotes both a buy and a sell price to earn the bid ask spread. Learn how market makers manage inventory, adverse selection and risk.
- Related: [Overfitting and Curve Fitting](https://learn.tradelabsai.com/research/overfitting-and-curve-fitting/): Overfitting means a strategy fits noise instead of a real pattern. Learn the warning signs, why it happens, how to measure it and practical ways to avoid it.
