# Monitoring and Logging Systems

> The tools behind trading observability: structured logs, metrics, dashboards, tracing and alerting with Prometheus, Grafana and log stacks, and how to set them up.

Source: https://learn.tradelabsai.com/infrastructure/monitoring-and-logging-systems/  
Track: Trading Infrastructure · Level: Advanced · Updated: 2026-10-03  
Publisher: TradeLabs AI (https://tradelabsai.com). Education, not financial advice.  
Cite as: TradeLabs Learn, "Monitoring and Logging Systems", https://learn.tradelabsai.com/infrastructure/monitoring-and-logging-systems/

Observability is the ability to understand what a system is doing from the outside, using the data it emits. For trading infrastructure, that data comes in three main forms: logs that record events, metrics that measure numbers over time and traces that follow a request across components. Together they answer the questions that matter during a busy session or an incident: is everything healthy, what changed and where did it go wrong? [Monitoring Positions, P&L and Risk](https://learn.tradelabsai.com/algo-trading/live-monitoring/) covers what to watch; this lesson covers the tools that make it possible.

## The three pillars

| Pillar | What it is | Trading examples |
|---|---|---|
| Logs | Timestamped records of events | Order sent, fill received, error raised |
| Metrics | Numeric measurements sampled over time | Message rate, latency percentiles, P&L, CPU use |
| Traces | The path of one request through services | A signal's journey from feed to order to fill |

## Structured logging

Write logs as structured records, such as JSON, rather than free text. Structured logs can be searched and aggregated reliably.

```python
import json, logging, time
log = logging.getLogger("bot")

def event(kind, **fields):
    log.info(json.dumps({"ts": time.time(), "event": kind, **fields}))

event("order_sent", client_order_id="s1-0007", symbol="ETHUSDT", side="buy", qty=0.5, px=3120.5)
```

Include consistent fields: timestamp, component, event type, symbol, order IDs and severity. Never log secrets or full API keys. See [Logging, Audit Trails and Incident Response](https://learn.tradelabsai.com/algo-trading/audit-trails/).

## Common tools

| Purpose | Tools |
|---|---|
| Metrics collection and alert rules | Prometheus |
| Dashboards | Grafana |
| Log storage and search | Loki, Elasticsearch or OpenSearch, cloud logging services |
| Tracing | OpenTelemetry with Jaeger or Tempo |
| Alert routing | Alertmanager, PagerDuty, Opsgenie, chat webhooks. See [Alerts and Webhooks](https://learn.tradelabsai.com/programming/alerts-and-webhooks/) |
| Uptime checks | External ping and HTTP monitoring services |

A small setup might be Prometheus and Grafana on one server plus log files with rotation. Larger firms run full observability platforms.

## Metrics worth exporting

| Category | Metrics |
|---|---|
| System | CPU, memory, disk usage, network errors |
| Feed | Messages per second, gaps, data age per symbol |
| Latency | Tick to trade and order round trip percentiles. See [Exchange vs Receive Timestamps and Latency Measurement](https://learn.tradelabsai.com/programming/latency-measurement/) |
| Orders | Sent, filled, rejected, cancelled per minute |
| Risk | Position, exposure, daily P&L, limit utilisation |
| Process | Heartbeat, restarts, queue depth |

**Example: A dashboard that catches a leak**
A Grafana dashboard shows a bot's memory use rising steadily from 300 MB after startup to 1.8 GB over three days, while message rates stay flat. An alert rule fires when memory passes 1.5 GB. Investigation finds the bot stores every trade it receives in a list that is never trimmed. Without the metric, the bot would have been killed by the operating system on the fourth day, probably in the middle of a session. A rolling buffer that keeps only the last hour fixes it. See [Alerts, Error Handling and Reconnection](https://learn.tradelabsai.com/algo-trading/error-handling/).

## Alerting rules

Good alert rules are based on symptoms that matter: data stale for more than 30 seconds, order rejects above a threshold, heartbeat missing, daily loss near the limit, disk above 85%. Route critical alerts to a phone and send lesser ones to a chat channel. Review alert history to remove noise. See [Monitoring Positions, P&L and Risk](https://learn.tradelabsai.com/algo-trading/live-monitoring/).

## Log retention and cost

Logs can be large. Keep detailed logs for a set period, archive compressed copies longer and delete per a written policy that meets any regulatory record keeping requirements. Sampling high volume debug logs reduces cost. See [Record Keeping for Traders](https://learn.tradelabsai.com/industry/record-keeping-for-traders/).

## Clocks matter

Correlating logs and metrics across servers requires synchronised clocks; otherwise events appear in the wrong order. See [Clock Synchronization and PTP](https://learn.tradelabsai.com/infrastructure/clock-synchronization-and-ptp/).

## Frequently asked questions

### What is the difference between logs and metrics?

Logs record individual events with details; metrics are numbers measured over time, such as rates and latencies, that are efficient to graph and alert on.

### What tools do traders use for monitoring?

Common choices are Prometheus for metrics, Grafana for dashboards, a log store such as Loki or Elasticsearch and an alert router that reaches phones and chat.

### What should a trading system log?

Signals, orders, acknowledgements, fills, rejects, risk checks, configuration changes, errors and system events, in structured form with consistent IDs.

Next, learn how to keep systems running through failures in [Fault Tolerance, High Availability and Redundancy](https://learn.tradelabsai.com/infrastructure/high-availability/).

## Continue learning

- Next lesson: [Fault Tolerance, High Availability and Redundancy](https://learn.tradelabsai.com/infrastructure/high-availability/)
- Previous lesson: [Redis and PostgreSQL for Trading](https://learn.tradelabsai.com/infrastructure/redis-and-postgresql-for-trading/)
- Related: [Redis and PostgreSQL for Trading](https://learn.tradelabsai.com/infrastructure/redis-and-postgresql-for-trading/): How Redis and PostgreSQL work together in trading systems: Redis for live prices, caches and state, PostgreSQL for durable orders, fills and history.
- Related: [Monitoring Positions, P&L and Risk](https://learn.tradelabsai.com/algo-trading/live-monitoring/): Running algorithms need constant monitoring. Learn the key health, trading and risk metrics to track, how to design useful alerts and how to avoid alert fatigue.
- Related: [Logging, Audit Trails and Incident Response](https://learn.tradelabsai.com/algo-trading/audit-trails/): An audit trail records every signal, order, change and fill so trading can be reconstructed later. Learn what to log, the regulatory rules and how it helps traders.
- Related: [Alerts and Webhooks](https://learn.tradelabsai.com/programming/alerts-and-webhooks/): How price alerts and webhooks work, how to send TradingView alerts to your own server or a chat channel, and how to secure webhook endpoints against fake signals.
- Related: [Fault Tolerance, High Availability and Redundancy](https://learn.tradelabsai.com/infrastructure/high-availability/): High availability keeps trading systems running through hardware, network and software failures. Learn redundancy, failover, avoiding split brain and testing it.
- Related: [Exchange vs Receive Timestamps and Latency Measurement](https://learn.tradelabsai.com/programming/latency-measurement/): How to measure latency in a trading system: where to timestamp, tick to trade and order round trip, percentiles instead of averages and how to find bottlenecks.
