TradeLabs AILearn

Fault Tolerance, High Availability and Redundancy

High availability keeps trading systems running through hardware, network and software failures. Learn redundancy, failover, avoiding split brain and testing it.

Advanced3 min readUpdated 3 Oct 2026
Markdown
Lesson 9 of 16

High availability (HA) means designing a system so that it keeps working when individual parts fail. Servers die, network links drop, data centres lose power and software crashes. An HA design has no single point of failure: for every critical component there is a backup ready to take over, ideally automatically and quickly. For trading systems, HA has a special twist. Two copies of a strategy both trading at once can be worse than none trading at all, so failover must be designed as carefully as redundancy.

Availability in numbers#

AvailabilityDowntime per year (approx.)
99%3.65 days
99.9%8.8 hours
99.99%53 minutes
99.999%5.3 minutes

What matters most in trading is availability during market hours, especially in volatile periods, when failures are both more likely and more costly.

Single points of failure#

ComponentRedundancy option
ServerStandby server, ideally in another zone or location
Network linkTwo providers or paths
Market data feedA and B lines, a backup data source. See Sequence Numbers, Dropped Packets and Out-of-Order Messages
Broker connectionA second session or a backup route
DatabaseReplica with automatic or manual failover. See Redis and PostgreSQL for Trading
PowerData centre redundancy; at home, a battery backup
PeopleMore than one person who can operate and stop the system

Active and passive setups#

SetupHow it worksTrade offs
Active passiveOne instance trades; a standby waits and takes over on failureSimple and safe, but failover takes time
Active active (for stateless services)Several instances share workGood for data and APIs; dangerous for order sending without coordination
Hot standbyStandby has live data and state, ready instantlyFast failover, more complexity
Cold standbyStandby must be started and loadedCheap, slower recovery

The split brain problem#

If the primary and standby lose contact with each other, each may believe the other has failed and both start trading. That doubles orders and positions. Defences:

  • A lock or lease held in a shared system (such as a database or coordination service) that only one instance can hold.
  • Fencing: the old primary is cut off from trading before the new one starts.
  • Broker side limits that would block a doubled position. See Risk Controls and Kill Switches.

State and failover#

The standby must know positions, open orders and strategy state. Options include sharing state through a replicated database, rebuilding state from the broker and recorded events at failover, or both. The broker's records are the final authority. See Trade Accounting and Reconciliation.

Testing HA#

Untested failover often fails. Run planned drills: stop the primary process, cut its network, kill the database primary and measure how long recovery takes and whether any orders were duplicated or lost. Chaos testing tools automate such failures in larger environments. See Failover, Backups and Disaster Recovery.

HA for retail traders#

A full HA setup is rarely worth it for one retail bot. Practical steps give most of the benefit:

  1. Run on a reliable cloud server with automatic restart. See VPS, Cloud and Bare-Metal Servers.
  2. Use broker side stops on every position.
  3. Monitor externally, so you know within a minute if the bot is down. See Monitoring Positions, P&L and Risk.
  4. Keep a documented manual fallback, such as the broker's mobile app.

Frequently asked questions#

What is high availability in trading?#

A design approach where systems keep running through component failures, using redundancy and failover with no single point of failure.

What is split brain?#

A failure where two instances both believe they are the active one, which in trading can cause duplicate orders and doubled positions.

Do retail traders need high availability?#

Usually not a full setup; a reliable server, automatic restarts, external monitoring and broker side stops cover most risks.

Next, learn why firms put servers inside exchange data centres in Co-Location.

Check your understanding

3 quick questions on this lesson. Get them all right to finish it.

Turn on JavaScript to take the quiz.

Finished this lesson?Sign in to save your progress across devices.
Next lessonCo-LocationCo location places trading servers inside or beside an exchange's data centre to cut latency. Learn how it works, what it costs, fairness rules and who needs it.

Mentioned in