# CPU Affinity, NUMA and Cache Optimization

> CPU affinity pins trading threads to specific cores so they are not interrupted or moved. Learn core isolation, NUMA, interrupts and how these cut latency spikes.

Source: https://learn.tradelabsai.com/infrastructure/cpu-affinity/  
Track: Trading Infrastructure · Level: Advanced · Updated: 2026-10-03  
Publisher: TradeLabs AI (https://tradelabsai.com). Education, not financial advice.  
Cite as: TradeLabs Learn, "CPU Affinity, NUMA and Cache Optimization", https://learn.tradelabsai.com/infrastructure/cpu-affinity/

Operating systems normally move programs between CPU cores freely and interrupt them whenever other work needs attention. That is great for general computing but bad for trading systems that need consistent reaction times. CPU affinity means pinning a thread or process to specific cores. Combined with isolating those cores from other work, it keeps a critical trading thread running undisturbed on its own core, with its data warm in that core's cache. The result is fewer latency spikes, which often matter more than the average speed.

## Why moving between cores hurts

| Effect | Cost |
|---|---|
| Context switches | Saving and restoring thread state when another task runs |
| Cache misses | A thread moved to a new core loses its warm cache |
| Interrupts | Hardware and timer interrupts pause the thread |
| Scheduling delays | A ready thread waits for a free core |

Each can add microseconds to milliseconds of delay at unpredictable times.

## Techniques

| Technique | What it does |
|---|---|
| Thread pinning | Bind a thread to one core (taskset, pthread_setaffinity_np, os.sched_setaffinity in Python) |
| Core isolation | Remove cores from the general scheduler (isolcpus or cpusets) |
| Interrupt steering | Move hardware interrupts away from trading cores (IRQ affinity) |
| Tickless cores | Reduce timer interrupts on isolated cores (nohz_full) |
| Disable power saving | Prevent cores from slowing down or sleeping (C states, frequency scaling) |
| Disable hyperthreading on critical cores | Avoid sharing core resources with another thread |

```
# Pin a process to core 3 on Linux
taskset -c 3 ./feed_handler

# Python: pin the current process to cores 2 and 3
import os
os.sched_setaffinity(0, {2, 3})
```

## NUMA awareness

Servers with multiple CPU sockets have NUMA (non uniform memory access): each socket has its own local memory, and reaching the other socket's memory is slower. Place trading threads, their memory and the network card on the same socket. The network card is attached to one socket; threads reading from it should run there. Tools like numactl help.

**Example: Taming the 99th percentile**
A firm's strategy thread shows a median reaction time of 4 microseconds but a 99.9th percentile of 120 microseconds. Tracing shows the slow cases coincide with the thread being moved between cores and with timer interrupts. After isolating cores 4 to 7, pinning the feed handler to core 4 and the strategy to core 5, moving network interrupts to core 2 and disabling deep power saving states, the median improves slightly to 3.5 microseconds, but the 99.9th percentile drops to 9 microseconds. The figures are illustrative; the pattern of tail improvement is typical. See [Exchange vs Receive Timestamps and Latency Measurement](https://learn.tradelabsai.com/programming/latency-measurement/).

## Typical core layout for a trading server

| Cores | Role |
|---|---|
| 0 and 1 | Operating system, logging, monitoring |
| 2 | Network interrupts |
| 3 | Feed handler (busy polling). See [Kernel Bypass and Low-Latency Networking](https://learn.tradelabsai.com/infrastructure/kernel-bypass/) |
| 4 | Strategy |
| 5 | Order gateway |
| Others | Non critical services, research |

Threads communicate through lock free queues in shared memory. See [Lock-Free Programming and Ring Buffers](https://learn.tradelabsai.com/infrastructure/lock-free-programming/).

## Costs and cautions

- **Dedicated cores burn power** at 100% when busy polling.
- **Misconfiguration can starve** the system of cores for essential tasks.
- **Cloud machines** may not expose real cores or allow full isolation; dedicated or bare metal servers give more control. See [VPS, Cloud and Bare-Metal Servers](https://learn.tradelabsai.com/infrastructure/vps-cloud-and-bare-metal-servers/).
- **Measure before and after;** tuning without measurement is guesswork.

## Who needs it

Latency sensitive firms use these techniques routinely. A retail bot reacting on second or minute timescales gains nothing meaningful from them. See [Trading Infrastructure Explained](https://learn.tradelabsai.com/infrastructure/trading-infrastructure-explained/).

## Frequently asked questions

### What is CPU affinity?

Binding a process or thread to specific CPU cores so the operating system does not move it, reducing cache misses and scheduling delays.

### What does isolcpus do?

It removes chosen cores from the Linux scheduler's general pool, so only processes explicitly pinned to them run there.

### Why does NUMA matter for trading?

Accessing memory or devices attached to the other CPU socket is slower, so placing threads, memory and network cards on the same socket reduces latency.

Next, learn how threads share data without locks in [Lock-Free Programming and Ring Buffers](https://learn.tradelabsai.com/infrastructure/lock-free-programming/).

## Continue learning

- Next lesson: [Lock-Free Programming and Ring Buffers](https://learn.tradelabsai.com/infrastructure/lock-free-programming/)
- Previous lesson: [FPGAs and Hardware Acceleration](https://learn.tradelabsai.com/infrastructure/fpgas-and-hardware-acceleration/)
- Related: [FPGAs and Hardware Acceleration](https://learn.tradelabsai.com/infrastructure/fpgas-and-hardware-acceleration/): FPGAs run trading logic in custom hardware for nanosecond reaction times. Learn what FPGAs are, what firms put on them, the costs, the limits and alternatives.
- Related: [Kernel Bypass and Low-Latency Networking](https://learn.tradelabsai.com/infrastructure/kernel-bypass/): Kernel bypass lets trading software read network packets straight from the network card, skipping the operating system. Learn how it works and the trade offs.
- Related: [Lock-Free Programming and Ring Buffers](https://learn.tradelabsai.com/infrastructure/lock-free-programming/): Lock free data structures let trading threads share data without waiting on locks. Learn ring buffers, atomics, the LMAX Disruptor pattern and the pitfalls involved.
- Related: [Exchange vs Receive Timestamps and Latency Measurement](https://learn.tradelabsai.com/programming/latency-measurement/): How to measure latency in a trading system: where to timestamp, tick to trade and order round trip, percentiles instead of averages and how to find bottlenecks.
- Related: [Linux for Traders](https://learn.tradelabsai.com/infrastructure/linux-for-traders/): The Linux skills traders need to run bots on servers: the command line, files, processes, services with systemd, logs, scheduling, SSH security and updates.
- Related: [Networking for Traders](https://learn.tradelabsai.com/infrastructure/networking-for-traders/): The networking basics behind trading: latency and distance, TCP versus UDP, multicast market data, packet loss, jitter and practical ways to improve connectivity.
