TradeLabs AILearn

Reinforcement Learning

Reinforcement learning trains agents to act by rewarding good outcomes. Learn how it applies to trading and execution, how rewards are designed and why it is hard.

Advanced3 min readUpdated 3 Oct 2026
Markdown
Read firstOnline Learning
Lesson 10 of 10

Reinforcement learning (RL) is a branch of machine learning in which an agent learns to make decisions by acting in an environment and receiving rewards or penalties. It famously mastered games such as Go and chess. Trading looks like a natural fit: an agent observes the market, chooses to buy, sell or hold, and is rewarded with profit. In practice, RL in trading is difficult, because markets are noisy, change over time and cannot be replayed endlessly with your actions affecting them realistically. Its most promising uses are narrow, well defined problems such as order execution and market making.

The RL framework#

ElementMeaningTrading example
AgentThe decision makerThe trading algorithm
EnvironmentWhat the agent interacts withThe market or a simulator
StateWhat the agent observesPrices, position, time remaining, order book
ActionWhat the agent can doBuy, sell, hold, place a limit order at a price
RewardFeedback after actionsProfit, minus costs and risk penalties
PolicyThe agent's strategyA mapping from states to actions

Main approaches#

ApproachIdeaExamples
Value basedLearn how good each action is in each stateQ learning, deep Q networks
Policy basedLearn the policy directlyPolicy gradients, PPO
Actor criticCombine bothA2C, SAC
Model basedLearn a model of the environment and planUsed where simulators are reliable

Where RL fits best in trading#

ProblemWhy RL suits it
Optimal executionClear goal (minimise cost), limited horizon, actions affect outcomes. See Optimal Execution and the Almgren-Chriss Model
Market makingRepeated decisions on quote placement and inventory. See Market Making
Hedging derivativesBalancing hedging costs and risk, sometimes called deep hedging. See Delta Hedging
Portfolio rebalancingTrading off tracking error and costs. See Rebalancing

Designing the reward#

The reward shapes everything the agent learns. Rewarding raw profit encourages excessive risk. Better rewards include risk adjusted returns, penalties for large positions or drawdowns, and transaction costs. Poorly designed rewards produce agents that exploit flaws in the simulator rather than learning real skill.

Why RL is hard in trading#

  • Limited data: you cannot replay the market millions of times like a game.
  • Simulators are imperfect: historical replay ignores how your orders would have moved prices. See Market Impact.
  • Non stationarity: the environment changes, so learned policies go stale. See Structural Breaks and Regime Changes.
  • Noisy rewards: profit is dominated by randomness, making learning slow and unstable.
  • Overfitting: agents easily memorise the training period. See Overfitting and Curve Fitting.
  • Safety: exploration with real money is costly. See Risk Controls and Kill Switches.

Practical advice#

  1. Start with a narrow problem with clear rewards, such as execution.
  2. Build a realistic simulator and validate it against real outcomes.
  3. Compare with simple baselines, such as TWAP or a rule based policy.
  4. Constrain actions to safe ranges.
  5. Test out of sample across different market conditions. See Walk-Forward Validation and Preventing Overfitting.

Frequently asked questions#

What is reinforcement learning in trading?#

A machine learning method where an agent learns trading or execution decisions by acting in a market environment and receiving rewards such as risk adjusted profit.

Does reinforcement learning work for trading?#

It shows promise in narrow problems such as execution, market making and hedging; using it to predict and trade prices directly is much harder.

Why is reward design important?#

The agent optimises whatever reward it is given, so rewards must include costs and risk, or it will learn risky or simulator specific behaviour.

You have finished the Machine Learning track. Continue with how to measure performance in Measuring Returns and CAGR.

Check your understanding

3 quick questions on this lesson. Get them all right to finish it.

Turn on JavaScript to take the quiz.

Finished this lesson?Sign in to save your progress across devices.

Where this leads