# Alerts, Error Handling and Reconnection

> Trading systems face rejected orders, disconnects, bad data and partial fills. Learn how to classify errors, retry safely, use idempotent orders and fail closed.

Source: https://learn.tradelabsai.com/algo-trading/error-handling/  
Track: Algorithmic Trading · Level: Advanced · Updated: 2026-10-03  
Publisher: TradeLabs AI (https://tradelabsai.com). Education, not financial advice.  
Cite as: TradeLabs Learn, "Alerts, Error Handling and Reconnection", https://learn.tradelabsai.com/algo-trading/error-handling/

In trading systems, errors are not rare events; they are part of daily life. Orders get rejected, connections drop, data arrives late or wrong, fills come in pieces and APIs return unexpected responses. The question is not whether errors happen but whether the system handles them safely. Good error handling follows one principle above all: when in doubt, do not trade. A system that stops and asks for help is far better than one that keeps guessing with real money.

## Common error types

| Error | Example | Safe response |
|---|---|---|
| Order rejected | Insufficient margin, invalid price, market closed | Log, alert, do not blindly resend |
| Connection lost | Network or broker outage | Stop new orders, reconnect, reconcile before resuming |
| Timeout | No response to an order | Query order status before retrying |
| Partial fill | Only part of an order filled | Track remaining quantity accurately |
| Bad data | Price spike, zero price, out of order ticks | Filter, validate, pause if persistent. See [Cleaning Market Data](https://learn.tradelabsai.com/programming/cleaning-market-data/) |
| Stale data | Feed stops updating silently | Detect and halt trading on that symbol |
| Rate limited | API refuses requests | Back off and slow down |
| Unexpected state | Position differs from expectations | Halt and reconcile. See [Trade Accounting and Reconciliation](https://learn.tradelabsai.com/industry/trade-reconciliation/) |

## Classifying errors

| Category | Meaning | Action |
|---|---|---|
| Transient | Likely to succeed if retried, such as a brief timeout | Retry with backoff |
| Permanent | Will fail again, such as an invalid symbol | Do not retry; alert |
| Unknown outcome | Not sure whether the action happened | Check state before doing anything else |
| Critical | Threatens capital or system integrity | Trigger the kill switch. See [Risk Controls and Kill Switches](https://learn.tradelabsai.com/algo-trading/risk-controls-and-kill-switches/) |

## The unknown outcome problem

The most dangerous error is when you do not know whether an order went through, for example after a timeout. Resending could double the position; not resending could miss the trade. The safe approach:

1. **Use a unique client order ID** for every order.
2. **Query the broker** for that ID before resending.
3. **Resend only if** the broker confirms the order does not exist.
4. **If the broker is unreachable,** stop sending new orders until state is confirmed.

This makes order submission idempotent: sending the same request twice has the same effect as sending it once.

**Example: Retrying with backoff**
A bot's request for account data fails with a timeout. Instead of retrying immediately in a tight loop, which could trigger rate limits, it waits 1 second, then 2, then 4, then 8, adding a small random delay each time. After five failures, it stops, raises a critical alert and pauses trading. Total waiting time before giving up is about 31 seconds, enough to ride out brief network glitches without hammering the broker's servers.

## Fail closed, not open

When a check cannot be completed, the safe default is to block the trade. If the risk check service is down, no orders should pass. If the price feed is stale, no new signals should act. Failing open, letting trades through when checks fail, is how small problems become large losses.

## Logging errors well

Every error log should include the time, component, error type, the order or symbol involved, the system state and what action was taken. This makes later diagnosis possible. See [Logging, Audit Trails and Incident Response](https://learn.tradelabsai.com/algo-trading/audit-trails/).

## Testing error paths

Error handling code runs rarely, so bugs hide there. Test it deliberately: disconnect the network, inject bad prices, simulate rejections and partial fills, and kill the process mid order. Confirm the system recovers correctly every time. See [Failover, Backups and Disaster Recovery](https://learn.tradelabsai.com/algo-trading/disaster-recovery/).

## Frequently asked questions

### How should a trading bot handle a rejected order?

Log the rejection reason, alert a human if needed and avoid blindly resending; fix the cause, such as insufficient margin or an invalid price, first.

### What does idempotent order submission mean?

Designing order requests, usually with unique client order IDs, so that sending the same request twice cannot create two orders.

### What does fail closed mean?

When a safety check cannot run, the system blocks trading rather than allowing it, so failures stop activity instead of letting risky orders through.

Next, learn how to prepare for bigger failures in [Failover, Backups and Disaster Recovery](https://learn.tradelabsai.com/algo-trading/disaster-recovery/).

## Continue learning

- Next lesson: [Failover, Backups and Disaster Recovery](https://learn.tradelabsai.com/algo-trading/disaster-recovery/)
- Previous lesson: [Monitoring Positions, P&L and Risk](https://learn.tradelabsai.com/algo-trading/live-monitoring/)
- Related: [Monitoring Positions, P&L and Risk](https://learn.tradelabsai.com/algo-trading/live-monitoring/): Running algorithms need constant monitoring. Learn the key health, trading and risk metrics to track, how to design useful alerts and how to avoid alert fatigue.
- Related: [Risk Controls and Kill Switches](https://learn.tradelabsai.com/algo-trading/risk-controls-and-kill-switches/): Pre trade risk checks and kill switches stop a trading algorithm before a bug becomes a disaster. Learn the essential limits, how to layer them and how to test them.
- Related: [Failover, Backups and Disaster Recovery](https://learn.tradelabsai.com/algo-trading/disaster-recovery/): Power cuts, server crashes and broker outages happen. Learn how traders and trading systems plan for disasters, with backups, failover and tested recovery steps.
- Related: [Automated vs Semi-Automated Trading](https://learn.tradelabsai.com/algo-trading/automated-trading/): Automated trading systems place and manage orders without manual input. Learn levels of automation, platforms, what to automate first and the safeguards.
- Related: [Trade Accounting and Reconciliation](https://learn.tradelabsai.com/industry/trade-reconciliation/): Reconciliation checks that internal records of trades, positions and cash match brokers, custodians and clearing houses. Learn the process, common breaks and fixes.
- Related: [Sequence Numbers, Dropped Packets and Out-of-Order Messages](https://learn.tradelabsai.com/programming/sequence-numbers/): Sequence numbers let trading systems detect lost, duplicated or out of order messages. Learn how gap detection, recovery and duplicate handling work in practice.
