Fault Tolerance, High Availability and Redundancy
High availability keeps trading systems running through hardware, network and software failures. Learn redundancy, failover, avoiding split brain and testing it.
High availability (HA) means designing a system so that it keeps working when individual parts fail. Servers die, network links drop, data centres lose power and software crashes. An HA design has no single point of failure: for every critical component there is a backup ready to take over, ideally automatically and quickly. For trading systems, HA has a special twist. Two copies of a strategy both trading at once can be worse than none trading at all, so failover must be designed as carefully as redundancy.
Availability in numbers#
| Availability | Downtime per year (approx.) |
|---|---|
| 99% | 3.65 days |
| 99.9% | 8.8 hours |
| 99.99% | 53 minutes |
| 99.999% | 5.3 minutes |
What matters most in trading is availability during market hours, especially in volatile periods, when failures are both more likely and more costly.
Single points of failure#
| Component | Redundancy option |
|---|---|
| Server | Standby server, ideally in another zone or location |
| Network link | Two providers or paths |
| Market data feed | A and B lines, a backup data source. See Sequence Numbers, Dropped Packets and Out-of-Order Messages |
| Broker connection | A second session or a backup route |
| Database | Replica with automatic or manual failover. See Redis and PostgreSQL for Trading |
| Power | Data centre redundancy; at home, a battery backup |
| People | More than one person who can operate and stop the system |
Active and passive setups#
| Setup | How it works | Trade offs |
|---|---|---|
| Active passive | One instance trades; a standby waits and takes over on failure | Simple and safe, but failover takes time |
| Active active (for stateless services) | Several instances share work | Good for data and APIs; dangerous for order sending without coordination |
| Hot standby | Standby has live data and state, ready instantly | Fast failover, more complexity |
| Cold standby | Standby must be started and loaded | Cheap, slower recovery |
The split brain problem#
If the primary and standby lose contact with each other, each may believe the other has failed and both start trading. That doubles orders and positions. Defences:
- A lock or lease held in a shared system (such as a database or coordination service) that only one instance can hold.
- Fencing: the old primary is cut off from trading before the new one starts.
- Broker side limits that would block a doubled position. See Risk Controls and Kill Switches.
State and failover#
The standby must know positions, open orders and strategy state. Options include sharing state through a replicated database, rebuilding state from the broker and recorded events at failover, or both. The broker's records are the final authority. See Trade Accounting and Reconciliation.
Testing HA#
Untested failover often fails. Run planned drills: stop the primary process, cut its network, kill the database primary and measure how long recovery takes and whether any orders were duplicated or lost. Chaos testing tools automate such failures in larger environments. See Failover, Backups and Disaster Recovery.
HA for retail traders#
A full HA setup is rarely worth it for one retail bot. Practical steps give most of the benefit:
- Run on a reliable cloud server with automatic restart. See VPS, Cloud and Bare-Metal Servers.
- Use broker side stops on every position.
- Monitor externally, so you know within a minute if the bot is down. See Monitoring Positions, P&L and Risk.
- Keep a documented manual fallback, such as the broker's mobile app.
Frequently asked questions#
What is high availability in trading?#
A design approach where systems keep running through component failures, using redundancy and failover with no single point of failure.
What is split brain?#
A failure where two instances both believe they are the active one, which in trading can cause duplicate orders and doubled positions.
Do retail traders need high availability?#
Usually not a full setup; a reliable server, automatic restarts, external monitoring and broker side stops cover most risks.
Next, learn why firms put servers inside exchange data centres in Co-Location.
3 quick questions on this lesson. Get them all right to finish it.
Turn on JavaScript to take the quiz.
Mentioned in
- Trading Infrastructure ExplainedTrading Infrastructure
- VPS, Cloud and Bare-Metal ServersTrading Infrastructure