TradeLabs AILearn

Monitoring and Logging Systems

The tools behind trading observability: structured logs, metrics, dashboards, tracing and alerting with Prometheus, Grafana and log stacks, and how to set them up.

Advanced3 min readUpdated 3 Oct 2026
Markdown
Lesson 8 of 16

Observability is the ability to understand what a system is doing from the outside, using the data it emits. For trading infrastructure, that data comes in three main forms: logs that record events, metrics that measure numbers over time and traces that follow a request across components. Together they answer the questions that matter during a busy session or an incident: is everything healthy, what changed and where did it go wrong? Monitoring Positions, P&L and Risk covers what to watch; this lesson covers the tools that make it possible.

The three pillars#

PillarWhat it isTrading examples
LogsTimestamped records of eventsOrder sent, fill received, error raised
MetricsNumeric measurements sampled over timeMessage rate, latency percentiles, P&L, CPU use
TracesThe path of one request through servicesA signal's journey from feed to order to fill

Structured logging#

Write logs as structured records, such as JSON, rather than free text. Structured logs can be searched and aggregated reliably.

import json, logging, time
log = logging.getLogger("bot")

def event(kind, **fields):
    log.info(json.dumps({"ts": time.time(), "event": kind, **fields}))

event("order_sent", client_order_id="s1-0007", symbol="ETHUSDT", side="buy", qty=0.5, px=3120.5)

Include consistent fields: timestamp, component, event type, symbol, order IDs and severity. Never log secrets or full API keys. See Logging, Audit Trails and Incident Response.

Common tools#

PurposeTools
Metrics collection and alert rulesPrometheus
DashboardsGrafana
Log storage and searchLoki, Elasticsearch or OpenSearch, cloud logging services
TracingOpenTelemetry with Jaeger or Tempo
Alert routingAlertmanager, PagerDuty, Opsgenie, chat webhooks. See Alerts and Webhooks
Uptime checksExternal ping and HTTP monitoring services

A small setup might be Prometheus and Grafana on one server plus log files with rotation. Larger firms run full observability platforms.

Metrics worth exporting#

CategoryMetrics
SystemCPU, memory, disk usage, network errors
FeedMessages per second, gaps, data age per symbol
LatencyTick to trade and order round trip percentiles. See Exchange vs Receive Timestamps and Latency Measurement
OrdersSent, filled, rejected, cancelled per minute
RiskPosition, exposure, daily P&L, limit utilisation
ProcessHeartbeat, restarts, queue depth

Alerting rules#

Good alert rules are based on symptoms that matter: data stale for more than 30 seconds, order rejects above a threshold, heartbeat missing, daily loss near the limit, disk above 85%. Route critical alerts to a phone and send lesser ones to a chat channel. Review alert history to remove noise. See Monitoring Positions, P&L and Risk.

Log retention and cost#

Logs can be large. Keep detailed logs for a set period, archive compressed copies longer and delete per a written policy that meets any regulatory record keeping requirements. Sampling high volume debug logs reduces cost. See Record Keeping for Traders.

Clocks matter#

Correlating logs and metrics across servers requires synchronised clocks; otherwise events appear in the wrong order. See Clock Synchronization and PTP.

Frequently asked questions#

What is the difference between logs and metrics?#

Logs record individual events with details; metrics are numbers measured over time, such as rates and latencies, that are efficient to graph and alert on.

What tools do traders use for monitoring?#

Common choices are Prometheus for metrics, Grafana for dashboards, a log store such as Loki or Elasticsearch and an alert router that reaches phones and chat.

What should a trading system log?#

Signals, orders, acknowledgements, fills, rejects, risk checks, configuration changes, errors and system events, in structured form with consistent IDs.

Next, learn how to keep systems running through failures in Fault Tolerance, High Availability and Redundancy.

Check your understanding

3 quick questions on this lesson. Get them all right to finish it.

Turn on JavaScript to take the quiz.

Finished this lesson?Sign in to save your progress across devices.
Next lessonFault Tolerance, High Availability and RedundancyHigh availability keeps trading systems running through hardware, network and software failures. Learn redundancy, failover, avoiding split brain and testing it.

Mentioned in