TradeLabs AILearn

CPU Affinity, NUMA and Cache Optimization

CPU affinity pins trading threads to specific cores so they are not interrupted or moved. Learn core isolation, NUMA, interrupts and how these cut latency spikes.

Advanced3 min readUpdated 3 Oct 2026
Markdown
Lesson 13 of 16

Operating systems normally move programs between CPU cores freely and interrupt them whenever other work needs attention. That is great for general computing but bad for trading systems that need consistent reaction times. CPU affinity means pinning a thread or process to specific cores. Combined with isolating those cores from other work, it keeps a critical trading thread running undisturbed on its own core, with its data warm in that core's cache. The result is fewer latency spikes, which often matter more than the average speed.

Why moving between cores hurts#

EffectCost
Context switchesSaving and restoring thread state when another task runs
Cache missesA thread moved to a new core loses its warm cache
InterruptsHardware and timer interrupts pause the thread
Scheduling delaysA ready thread waits for a free core

Each can add microseconds to milliseconds of delay at unpredictable times.

Techniques#

TechniqueWhat it does
Thread pinningBind a thread to one core (taskset, pthread_setaffinity_np, os.sched_setaffinity in Python)
Core isolationRemove cores from the general scheduler (isolcpus or cpusets)
Interrupt steeringMove hardware interrupts away from trading cores (IRQ affinity)
Tickless coresReduce timer interrupts on isolated cores (nohz_full)
Disable power savingPrevent cores from slowing down or sleeping (C states, frequency scaling)
Disable hyperthreading on critical coresAvoid sharing core resources with another thread
# Pin a process to core 3 on Linux
taskset -c 3 ./feed_handler

# Python: pin the current process to cores 2 and 3
import os
os.sched_setaffinity(0, {2, 3})

NUMA awareness#

Servers with multiple CPU sockets have NUMA (non uniform memory access): each socket has its own local memory, and reaching the other socket's memory is slower. Place trading threads, their memory and the network card on the same socket. The network card is attached to one socket; threads reading from it should run there. Tools like numactl help.

Typical core layout for a trading server#

CoresRole
0 and 1Operating system, logging, monitoring
2Network interrupts
3Feed handler (busy polling). See Kernel Bypass and Low-Latency Networking
4Strategy
5Order gateway
OthersNon critical services, research

Threads communicate through lock free queues in shared memory. See Lock-Free Programming and Ring Buffers.

Costs and cautions#

  • Dedicated cores burn power at 100% when busy polling.
  • Misconfiguration can starve the system of cores for essential tasks.
  • Cloud machines may not expose real cores or allow full isolation; dedicated or bare metal servers give more control. See VPS, Cloud and Bare-Metal Servers.
  • Measure before and after; tuning without measurement is guesswork.

Who needs it#

Latency sensitive firms use these techniques routinely. A retail bot reacting on second or minute timescales gains nothing meaningful from them. See Trading Infrastructure Explained.

Frequently asked questions#

What is CPU affinity?#

Binding a process or thread to specific CPU cores so the operating system does not move it, reducing cache misses and scheduling delays.

What does isolcpus do?#

It removes chosen cores from the Linux scheduler's general pool, so only processes explicitly pinned to them run there.

Why does NUMA matter for trading?#

Accessing memory or devices attached to the other CPU socket is slower, so placing threads, memory and network cards on the same socket reduces latency.

Next, learn how threads share data without locks in Lock-Free Programming and Ring Buffers.

Check your understanding

3 quick questions on this lesson. Get them all right to finish it.

Turn on JavaScript to take the quiz.

Finished this lesson?Sign in to save your progress across devices.
Next lessonLock-Free Programming and Ring BuffersLock free data structures let trading threads share data without waiting on locks. Learn ring buffers, atomics, the LMAX Disruptor pattern and the pitfalls involved.

Mentioned in