# Fault Tolerance, High Availability and Redundancy

> High availability keeps trading systems running through hardware, network and software failures. Learn redundancy, failover, avoiding split brain and testing it.

Source: https://learn.tradelabsai.com/infrastructure/high-availability/  
Track: Trading Infrastructure · Level: Advanced · Updated: 2026-10-03  
Publisher: TradeLabs AI (https://tradelabsai.com). Education, not financial advice.  
Cite as: TradeLabs Learn, "Fault Tolerance, High Availability and Redundancy", https://learn.tradelabsai.com/infrastructure/high-availability/

High availability (HA) means designing a system so that it keeps working when individual parts fail. Servers die, network links drop, data centres lose power and software crashes. An HA design has no single point of failure: for every critical component there is a backup ready to take over, ideally automatically and quickly. For trading systems, HA has a special twist. Two copies of a strategy both trading at once can be worse than none trading at all, so failover must be designed as carefully as redundancy.

## Availability in numbers

| Availability | Downtime per year (approx.) |
|---|---|
| 99% | 3.65 days |
| 99.9% | 8.8 hours |
| 99.99% | 53 minutes |
| 99.999% | 5.3 minutes |

What matters most in trading is availability during market hours, especially in volatile periods, when failures are both more likely and more costly.

## Single points of failure

| Component | Redundancy option |
|---|---|
| Server | Standby server, ideally in another zone or location |
| Network link | Two providers or paths |
| Market data feed | A and B lines, a backup data source. See [Sequence Numbers, Dropped Packets and Out-of-Order Messages](https://learn.tradelabsai.com/programming/sequence-numbers/) |
| Broker connection | A second session or a backup route |
| Database | Replica with automatic or manual failover. See [Redis and PostgreSQL for Trading](https://learn.tradelabsai.com/infrastructure/redis-and-postgresql-for-trading/) |
| Power | Data centre redundancy; at home, a battery backup |
| People | More than one person who can operate and stop the system |

## Active and passive setups

| Setup | How it works | Trade offs |
|---|---|---|
| Active passive | One instance trades; a standby waits and takes over on failure | Simple and safe, but failover takes time |
| Active active (for stateless services) | Several instances share work | Good for data and APIs; dangerous for order sending without coordination |
| Hot standby | Standby has live data and state, ready instantly | Fast failover, more complexity |
| Cold standby | Standby must be started and loaded | Cheap, slower recovery |

## The split brain problem

If the primary and standby lose contact with each other, each may believe the other has failed and both start trading. That doubles orders and positions. Defences:

- **A lock or lease** held in a shared system (such as a database or coordination service) that only one instance can hold.
- **Fencing:** the old primary is cut off from trading before the new one starts.
- **Broker side limits** that would block a doubled position. See [Risk Controls and Kill Switches](https://learn.tradelabsai.com/algo-trading/risk-controls-and-kill-switches/).

**Example: A safe failover**
A bot runs on a primary server and holds a lease in a database that it renews every 5 seconds; the lease expires after 15 seconds without renewal. The primary's network fails at 10:15:00. Its lease expires at about 10:15:15. The standby, checking every 2 seconds, acquires the lease at around 10:15:16, reconciles positions and open orders with the broker, and resumes trading by 10:15:20. When the primary's network returns, it sees it no longer holds the lease and stays idle. Roughly 20 seconds of downtime, and never two active traders.

## State and failover

The standby must know positions, open orders and strategy state. Options include sharing state through a replicated database, rebuilding state from the broker and recorded events at failover, or both. The broker's records are the final authority. See [Trade Accounting and Reconciliation](https://learn.tradelabsai.com/industry/trade-reconciliation/).

## Testing HA

Untested failover often fails. Run planned drills: stop the primary process, cut its network, kill the database primary and measure how long recovery takes and whether any orders were duplicated or lost. Chaos testing tools automate such failures in larger environments. See [Failover, Backups and Disaster Recovery](https://learn.tradelabsai.com/algo-trading/disaster-recovery/).

## HA for retail traders

A full HA setup is rarely worth it for one retail bot. Practical steps give most of the benefit:

1. **Run on a reliable cloud server** with automatic restart. See [VPS, Cloud and Bare-Metal Servers](https://learn.tradelabsai.com/infrastructure/vps-cloud-and-bare-metal-servers/).
2. **Use broker side stops** on every position.
3. **Monitor externally,** so you know within a minute if the bot is down. See [Monitoring Positions, P&L and Risk](https://learn.tradelabsai.com/algo-trading/live-monitoring/).
4. **Keep a documented manual fallback,** such as the broker's mobile app.

## Frequently asked questions

### What is high availability in trading?

A design approach where systems keep running through component failures, using redundancy and failover with no single point of failure.

### What is split brain?

A failure where two instances both believe they are the active one, which in trading can cause duplicate orders and doubled positions.

### Do retail traders need high availability?

Usually not a full setup; a reliable server, automatic restarts, external monitoring and broker side stops cover most risks.

Next, learn why firms put servers inside exchange data centres in [Co-Location](https://learn.tradelabsai.com/infrastructure/co-location/).

## Continue learning

- Next lesson: [Co-Location](https://learn.tradelabsai.com/infrastructure/co-location/)
- Previous lesson: [Monitoring and Logging Systems](https://learn.tradelabsai.com/infrastructure/monitoring-and-logging-systems/)
- Related: [Monitoring and Logging Systems](https://learn.tradelabsai.com/infrastructure/monitoring-and-logging-systems/): The tools behind trading observability: structured logs, metrics, dashboards, tracing and alerting with Prometheus, Grafana and log stacks, and how to set them up.
- Related: [Failover, Backups and Disaster Recovery](https://learn.tradelabsai.com/algo-trading/disaster-recovery/): Power cuts, server crashes and broker outages happen. Learn how traders and trading systems plan for disasters, with backups, failover and tested recovery steps.
- Related: [Redis and PostgreSQL for Trading](https://learn.tradelabsai.com/infrastructure/redis-and-postgresql-for-trading/): How Redis and PostgreSQL work together in trading systems: Redis for live prices, caches and state, PostgreSQL for durable orders, fills and history.
- Related: [Docker and Kubernetes](https://learn.tradelabsai.com/infrastructure/docker-and-kubernetes/): How containers with Docker package trading bots and research environments, when Kubernetes helps, and when simpler setups are better for latency and reliability.
- Related: [Alerts, Error Handling and Reconnection](https://learn.tradelabsai.com/algo-trading/error-handling/): Trading systems face rejected orders, disconnects, bad data and partial fills. Learn how to classify errors, retry safely, use idempotent orders and fail closed.
