Reinforcement Learning
Reinforcement learning trains agents to act by rewarding good outcomes. Learn how it applies to trading and execution, how rewards are designed and why it is hard.
Reinforcement learning (RL) is a branch of machine learning in which an agent learns to make decisions by acting in an environment and receiving rewards or penalties. It famously mastered games such as Go and chess. Trading looks like a natural fit: an agent observes the market, chooses to buy, sell or hold, and is rewarded with profit. In practice, RL in trading is difficult, because markets are noisy, change over time and cannot be replayed endlessly with your actions affecting them realistically. Its most promising uses are narrow, well defined problems such as order execution and market making.
The RL framework#
| Element | Meaning | Trading example |
|---|---|---|
| Agent | The decision maker | The trading algorithm |
| Environment | What the agent interacts with | The market or a simulator |
| State | What the agent observes | Prices, position, time remaining, order book |
| Action | What the agent can do | Buy, sell, hold, place a limit order at a price |
| Reward | Feedback after actions | Profit, minus costs and risk penalties |
| Policy | The agent's strategy | A mapping from states to actions |
Main approaches#
| Approach | Idea | Examples |
|---|---|---|
| Value based | Learn how good each action is in each state | Q learning, deep Q networks |
| Policy based | Learn the policy directly | Policy gradients, PPO |
| Actor critic | Combine both | A2C, SAC |
| Model based | Learn a model of the environment and plan | Used where simulators are reliable |
Where RL fits best in trading#
| Problem | Why RL suits it |
|---|---|
| Optimal execution | Clear goal (minimise cost), limited horizon, actions affect outcomes. See Optimal Execution and the Almgren-Chriss Model |
| Market making | Repeated decisions on quote placement and inventory. See Market Making |
| Hedging derivatives | Balancing hedging costs and risk, sometimes called deep hedging. See Delta Hedging |
| Portfolio rebalancing | Trading off tracking error and costs. See Rebalancing |
Designing the reward#
The reward shapes everything the agent learns. Rewarding raw profit encourages excessive risk. Better rewards include risk adjusted returns, penalties for large positions or drawdowns, and transaction costs. Poorly designed rewards produce agents that exploit flaws in the simulator rather than learning real skill.
Why RL is hard in trading#
- Limited data: you cannot replay the market millions of times like a game.
- Simulators are imperfect: historical replay ignores how your orders would have moved prices. See Market Impact.
- Non stationarity: the environment changes, so learned policies go stale. See Structural Breaks and Regime Changes.
- Noisy rewards: profit is dominated by randomness, making learning slow and unstable.
- Overfitting: agents easily memorise the training period. See Overfitting and Curve Fitting.
- Safety: exploration with real money is costly. See Risk Controls and Kill Switches.
Practical advice#
- Start with a narrow problem with clear rewards, such as execution.
- Build a realistic simulator and validate it against real outcomes.
- Compare with simple baselines, such as TWAP or a rule based policy.
- Constrain actions to safe ranges.
- Test out of sample across different market conditions. See Walk-Forward Validation and Preventing Overfitting.
Frequently asked questions#
What is reinforcement learning in trading?#
A machine learning method where an agent learns trading or execution decisions by acting in a market environment and receiving rewards such as risk adjusted profit.
Does reinforcement learning work for trading?#
It shows promise in narrow problems such as execution, market making and hedging; using it to predict and trade prices directly is much harder.
Why is reward design important?#
The agent optimises whatever reward it is given, so rewards must include costs and risk, or it will learn risky or simulator specific behaviour.
You have finished the Machine Learning track. Continue with how to measure performance in Measuring Returns and CAGR.
3 quick questions on this lesson. Get them all right to finish it.
Turn on JavaScript to take the quiz.
Where this leads
- Measuring Returns and CAGRPortfolio and Performance