Monitoring and Logging Systems
The tools behind trading observability: structured logs, metrics, dashboards, tracing and alerting with Prometheus, Grafana and log stacks, and how to set them up.
Observability is the ability to understand what a system is doing from the outside, using the data it emits. For trading infrastructure, that data comes in three main forms: logs that record events, metrics that measure numbers over time and traces that follow a request across components. Together they answer the questions that matter during a busy session or an incident: is everything healthy, what changed and where did it go wrong? Monitoring Positions, P&L and Risk covers what to watch; this lesson covers the tools that make it possible.
The three pillars#
| Pillar | What it is | Trading examples |
|---|---|---|
| Logs | Timestamped records of events | Order sent, fill received, error raised |
| Metrics | Numeric measurements sampled over time | Message rate, latency percentiles, P&L, CPU use |
| Traces | The path of one request through services | A signal's journey from feed to order to fill |
Structured logging#
Write logs as structured records, such as JSON, rather than free text. Structured logs can be searched and aggregated reliably.
import json, logging, time
log = logging.getLogger("bot")
def event(kind, **fields):
log.info(json.dumps({"ts": time.time(), "event": kind, **fields}))
event("order_sent", client_order_id="s1-0007", symbol="ETHUSDT", side="buy", qty=0.5, px=3120.5)
Include consistent fields: timestamp, component, event type, symbol, order IDs and severity. Never log secrets or full API keys. See Logging, Audit Trails and Incident Response.
Common tools#
| Purpose | Tools |
|---|---|
| Metrics collection and alert rules | Prometheus |
| Dashboards | Grafana |
| Log storage and search | Loki, Elasticsearch or OpenSearch, cloud logging services |
| Tracing | OpenTelemetry with Jaeger or Tempo |
| Alert routing | Alertmanager, PagerDuty, Opsgenie, chat webhooks. See Alerts and Webhooks |
| Uptime checks | External ping and HTTP monitoring services |
A small setup might be Prometheus and Grafana on one server plus log files with rotation. Larger firms run full observability platforms.
Metrics worth exporting#
| Category | Metrics |
|---|---|
| System | CPU, memory, disk usage, network errors |
| Feed | Messages per second, gaps, data age per symbol |
| Latency | Tick to trade and order round trip percentiles. See Exchange vs Receive Timestamps and Latency Measurement |
| Orders | Sent, filled, rejected, cancelled per minute |
| Risk | Position, exposure, daily P&L, limit utilisation |
| Process | Heartbeat, restarts, queue depth |
Alerting rules#
Good alert rules are based on symptoms that matter: data stale for more than 30 seconds, order rejects above a threshold, heartbeat missing, daily loss near the limit, disk above 85%. Route critical alerts to a phone and send lesser ones to a chat channel. Review alert history to remove noise. See Monitoring Positions, P&L and Risk.
Log retention and cost#
Logs can be large. Keep detailed logs for a set period, archive compressed copies longer and delete per a written policy that meets any regulatory record keeping requirements. Sampling high volume debug logs reduces cost. See Record Keeping for Traders.
Clocks matter#
Correlating logs and metrics across servers requires synchronised clocks; otherwise events appear in the wrong order. See Clock Synchronization and PTP.
Frequently asked questions#
What is the difference between logs and metrics?#
Logs record individual events with details; metrics are numbers measured over time, such as rates and latencies, that are efficient to graph and alert on.
What tools do traders use for monitoring?#
Common choices are Prometheus for metrics, Grafana for dashboards, a log store such as Loki or Elasticsearch and an alert router that reaches phones and chat.
What should a trading system log?#
Signals, orders, acknowledgements, fills, rejects, risk checks, configuration changes, errors and system events, in structured form with consistent IDs.
Next, learn how to keep systems running through failures in Fault Tolerance, High Availability and Redundancy.
3 quick questions on this lesson. Get them all right to finish it.
Turn on JavaScript to take the quiz.
Mentioned in
- Trading Infrastructure ExplainedTrading Infrastructure
- VPS, Cloud and Bare-Metal ServersTrading Infrastructure
- Linux for TradersTrading Infrastructure
- Docker and KubernetesTrading Infrastructure
- Redis and PostgreSQL for TradingTrading Infrastructure
- Clock Synchronization and PTPTrading Infrastructure