# Lock-Free Programming and Ring Buffers

> Lock free data structures let trading threads share data without waiting on locks. Learn ring buffers, atomics, the LMAX Disruptor pattern and the pitfalls involved.

Source: https://learn.tradelabsai.com/infrastructure/lock-free-programming/  
Track: Trading Infrastructure · Level: Advanced · Updated: 2026-10-03  
Publisher: TradeLabs AI (https://tradelabsai.com). Education, not financial advice.  
Cite as: TradeLabs Learn, "Lock-Free Programming and Ring Buffers", https://learn.tradelabsai.com/infrastructure/lock-free-programming/

Low latency trading systems split work across threads: one reads market data, another runs the strategy, another sends orders. These threads must pass data to each other constantly. The traditional way to share data safely is a lock, which lets only one thread touch the data at a time. But locks make threads wait, and waiting creates unpredictable delays. Lock free programming uses special processor instructions and careful design so threads can exchange data without ever blocking each other.

## Why locks hurt latency

| Problem | Effect |
|---|---|
| Contention | Threads wait for each other to release the lock |
| Context switches | A waiting thread may be put to sleep and woken later |
| Priority inversion | A low priority thread holding a lock delays a high priority one |
| Unpredictability | Delays vary widely, hurting tail latency |

## Building blocks

| Concept | Meaning |
|---|---|
| Atomic operations | Instructions that complete as one indivisible step, such as atomic increment |
| Compare and swap (CAS) | Update a value only if it still equals what you expected |
| Memory ordering | Rules about when writes by one thread become visible to another |
| Cache lines | Memory is moved in blocks (commonly 64 bytes); sharing them between cores is costly |

## The ring buffer

The workhorse of low latency messaging is a fixed size circular buffer, often with one producer and one consumer (SPSC). The producer writes to the next slot and advances a write index; the consumer reads and advances a read index. Each thread writes only its own index, so with correct memory ordering, no lock is needed.

```
producer:                          consumer:
  wait until slot is free            wait until slot has data
  write item into slot               read item from slot
  publish new write index            publish new read index
```

## The LMAX Disruptor

In 2011, the London based exchange LMAX published the Disruptor, a lock free ring buffer design for Java that processed millions of events per second on a single thread with very low latency. Its ideas, such as preallocated ring buffers, sequence counters, avoiding false sharing and batching, have influenced many trading systems in Java, C++ and other languages. Martin Fowler's article on the LMAX architecture is a well known introduction.

**Example: False sharing**
Two threads update separate counters that happen to sit in the same 64 byte cache line. Although they never touch each other's counter, every write forces the cache line to bounce between the two cores, and a benchmark shows each thread managing only a fraction of its expected throughput. Padding each counter so it occupies its own cache line removes the conflict, and throughput rises several times. Nothing in the program's logic changed; only the memory layout did. This is called false sharing, and lock free designs must avoid it.

## Pitfalls

- **Correctness is hard:** subtle memory ordering bugs may appear only under rare timing.
- **The ABA problem:** a value changes from A to B and back to A, fooling a compare and swap.
- **Busy waiting** uses full CPU cores. See [CPU Affinity, NUMA and Cache Optimization](https://learn.tradelabsai.com/infrastructure/cpu-affinity/).
- **Multi producer designs** are much harder than single producer ones.
- **Testing:** use stress tests, sanitizers and, where possible, proven libraries.

## Practical advice

1. **Prefer single producer, single consumer** queues where possible; they are simpler and faster.
2. **Use well tested libraries** rather than writing your own.
3. **Preallocate memory** to avoid allocation delays.
4. **Measure** with realistic loads, focusing on tail latency. See [Exchange vs Receive Timestamps and Latency Measurement](https://learn.tradelabsai.com/programming/latency-measurement/).
5. **Do not bother for slow strategies:** standard queues are perfectly adequate for most bots. See [Message Queues](https://learn.tradelabsai.com/infrastructure/message-queues/).

## Where it is used

Exchange matching engines, market data handlers, order gateways and high frequency strategies commonly use lock free queues to pass events between pinned threads. See [Matching Engines](https://learn.tradelabsai.com/orders/matching-engines/).

## Frequently asked questions

### What is lock free programming?

A way of writing concurrent code where threads share data using atomic operations instead of locks, so no thread blocks waiting for another.

### What is the LMAX Disruptor?

A lock free ring buffer design published by the LMAX exchange in 2011 that achieved very high throughput and low latency, influencing many trading systems.

### Do I need lock free code for a trading bot?

Rarely. It matters for microsecond sensitive systems; most bots work fine with standard thread safe queues.

Next, learn why trading systems use compact binary messages in [Binary Protocols](https://learn.tradelabsai.com/infrastructure/binary-protocols/).

## Continue learning

- Next lesson: [Binary Protocols](https://learn.tradelabsai.com/infrastructure/binary-protocols/)
- Previous lesson: [CPU Affinity, NUMA and Cache Optimization](https://learn.tradelabsai.com/infrastructure/cpu-affinity/)
- Related: [CPU Affinity, NUMA and Cache Optimization](https://learn.tradelabsai.com/infrastructure/cpu-affinity/): CPU affinity pins trading threads to specific cores so they are not interrupted or moved. Learn core isolation, NUMA, interrupts and how these cut latency spikes.
- Related: [Message Queues](https://learn.tradelabsai.com/infrastructure/message-queues/): Message queues pass market data, signals and orders between trading components. Learn pub sub and queues, tools like Kafka, Redis and ZeroMQ, and design trade offs.
- Related: [Exchange vs Receive Timestamps and Latency Measurement](https://learn.tradelabsai.com/programming/latency-measurement/): How to measure latency in a trading system: where to timestamp, tick to trade and order round trip, percentiles instead of averages and how to find bottlenecks.
- Related: [Kernel Bypass and Low-Latency Networking](https://learn.tradelabsai.com/infrastructure/kernel-bypass/): Kernel bypass lets trading software read network packets straight from the network card, skipping the operating system. Learn how it works and the trade offs.
- Related: [Matching Engines](https://learn.tradelabsai.com/orders/matching-engines/): A matching engine is the system that pairs buy and sell orders on an exchange. Learn price time priority, pro rata matching, auctions and why it matters to you.
