# Failover, Backups and Disaster Recovery

> Power cuts, server crashes and broker outages happen. Learn how traders and trading systems plan for disasters, with backups, failover and tested recovery steps.

Source: https://learn.tradelabsai.com/algo-trading/disaster-recovery/  
Track: Algorithmic Trading · Level: Advanced · Updated: 2026-10-03  
Publisher: TradeLabs AI (https://tradelabsai.com). Education, not financial advice.  
Cite as: TradeLabs Learn, "Failover, Backups and Disaster Recovery", https://learn.tradelabsai.com/algo-trading/disaster-recovery/

Disaster recovery is the plan for what happens when something big breaks: a server dies, a data centre loses power, the internet goes down, a broker has an outage or a bad deployment corrupts the system. Trading makes these events especially costly because open positions keep moving while you are offline. A good recovery plan defines how quickly you need to be back, how you will get there and how you will manage open risk in the meantime. For a solo trader it can be a one page checklist; for a firm it is a tested, documented process.

## Key concepts

| Term | Meaning |
|---|---|
| Recovery time objective (RTO) | How quickly the system must be running again |
| Recovery point objective (RPO) | How much data loss is acceptable, measured in time |
| Failover | Switching to a backup system |
| Backup | A copy of data or configuration that can be restored |
| Runbook | Step by step instructions for a specific failure |

## Common disaster scenarios

| Scenario | Impact | Mitigation |
|---|---|---|
| Home internet or power outage | Cannot reach broker | Mobile hotspot, battery backup, broker phone desk, run bots on a remote server. See [VPS, Cloud and Bare-Metal Servers](https://learn.tradelabsai.com/infrastructure/vps-cloud-and-bare-metal-servers/) |
| Server crash | Bot stops; positions unmanaged | Automatic restart, standby server, broker side stops |
| Data centre outage | Whole system down | Second location or cloud region. See [Fault Tolerance, High Availability and Redundancy](https://learn.tradelabsai.com/infrastructure/high-availability/) |
| Broker outage | Cannot trade or exit | Accounts at a second broker, hedging with other instruments |
| Bad deployment | Wrong behaviour | Rollback plan, staged releases |
| Data corruption | Wrong positions or history | Regular backups, reconciliation with broker records |

## Broker side protection

The single most useful protection for many traders is placing stop orders at the broker rather than only inside a local program. If your system disappears, broker side stops still protect positions. Bracket and OCO orders keep exits alive even when you are offline. See [Bracket Orders](https://learn.tradelabsai.com/orders/bracket-orders/) and [OCO Orders](https://learn.tradelabsai.com/orders/oco-orders/). Note that stops do not protect against gaps. See [Price Gaps and How to Trade Them](https://learn.tradelabsai.com/chart-patterns/price-gaps-and-how-to-trade-them/).

## A recovery runbook

1. **Assess:** what failed and what positions and orders are open?
2. **Protect:** ensure open positions have protective orders; if unsure, reduce risk through any available channel, including the broker's phone desk.
3. **Restore:** bring up the backup system or restore from backup.
4. **Reconcile:** confirm positions and orders match the broker's records before trading resumes. See [Trade Accounting and Reconciliation](https://learn.tradelabsai.com/industry/trade-reconciliation/).
5. **Resume carefully:** restart with reduced size or in a monitoring mode.
6. **Review:** document what happened and improve the plan.

**Example: A home trader's outage**
A trader runs a futures bot on a home computer. A storm cuts power while the bot holds a long position of two contracts. Because the bot places a stop at the broker with every entry, the position is protected. The trader uses a phone's mobile app to confirm the stop is live. After the outage, the trader moves the bot to a cloud server for about the cost of a few dollars to tens of dollars a month and adds a heartbeat alert that sends a message if the bot stops reporting. See [Monitoring Positions, P&L and Risk](https://learn.tradelabsai.com/algo-trading/live-monitoring/).

## Backups

| What to back up | How often |
|---|---|
| Code | Every change, in version control |
| Configuration and parameters | Every change |
| Trade and order records | Daily or continuously |
| Historical data | Regularly, with checksums |
| Credentials | Stored securely, recoverable by authorised people only |

A backup that has never been restored is not proven to work. Test restores periodically.

## Testing the plan

Run drills: shut down the main server and fail over, restore a database from backup, practise reaching the broker by phone. Time each step and compare against your recovery time objective. Firms often run scheduled disaster recovery tests, and some regulations require them.

## Frequently asked questions

### What is disaster recovery in trading?

The plan and tools for restoring trading systems and protecting open positions after major failures such as outages, crashes or data loss.

### How can retail traders protect against outages?

Place stops at the broker, keep a mobile app and broker phone number ready, use backup internet and power, and run automated systems on a remote server.

### What are RTO and RPO?

Recovery time objective is how fast you must recover; recovery point objective is how much recent data you can afford to lose.

Next, learn how to keep complete records of every action in [Logging, Audit Trails and Incident Response](https://learn.tradelabsai.com/algo-trading/audit-trails/).

## Continue learning

- Next lesson: [Logging, Audit Trails and Incident Response](https://learn.tradelabsai.com/algo-trading/audit-trails/)
- Previous lesson: [Alerts, Error Handling and Reconnection](https://learn.tradelabsai.com/algo-trading/error-handling/)
- Related: [Alerts, Error Handling and Reconnection](https://learn.tradelabsai.com/algo-trading/error-handling/): Trading systems face rejected orders, disconnects, bad data and partial fills. Learn how to classify errors, retry safely, use idempotent orders and fail closed.
- Related: [Fault Tolerance, High Availability and Redundancy](https://learn.tradelabsai.com/infrastructure/high-availability/): High availability keeps trading systems running through hardware, network and software failures. Learn redundancy, failover, avoiding split brain and testing it.
- Related: [VPS, Cloud and Bare-Metal Servers](https://learn.tradelabsai.com/infrastructure/vps-cloud-and-bare-metal-servers/): Compare VPS, cloud and bare metal servers for running trading bots and systems. Learn costs, latency, reliability, choosing a region and basic server setup.
- Related: [Operational and Model Risk](https://learn.tradelabsai.com/portfolio/operational-and-model-risk/): Operational risk comes from failed processes, people and systems; model risk from wrong or misused models. Learn real examples and the key controls.
- Related: [Trade Accounting and Reconciliation](https://learn.tradelabsai.com/industry/trade-reconciliation/): Reconciliation checks that internal records of trades, positions and cash match brokers, custodians and clearing houses. Learn the process, common breaks and fixes.
