# Kernel Bypass and Low-Latency Networking

> Kernel bypass lets trading software read network packets straight from the network card, skipping the operating system. Learn how it works and the trade offs.

Source: https://learn.tradelabsai.com/infrastructure/kernel-bypass/  
Track: Trading Infrastructure · Level: Advanced · Updated: 2026-10-03  
Publisher: TradeLabs AI (https://tradelabsai.com). Education, not financial advice.  
Cite as: TradeLabs Learn, "Kernel Bypass and Low-Latency Networking", https://learn.tradelabsai.com/infrastructure/kernel-bypass/

When a network packet arrives at a normal server, the operating system kernel handles it first: interrupts fire, the packet is copied into kernel memory, passed through the network stack and finally copied again into the application. Each step costs time and adds unpredictable delays. Kernel bypass techniques let the trading application read packets directly from the network card's memory, skipping most of that path. For high frequency firms, kernel bypass can cut network processing from several microseconds to around a microsecond or less, with far more consistent timing.

## The normal path versus kernel bypass

| Step | Normal kernel networking | Kernel bypass |
|---|---|---|
| Packet arrives | Network card raises an interrupt | Application polls the card directly |
| Processing | Kernel network stack (IP, TCP or UDP) | User space library or the application |
| Copies | Into kernel buffers, then into the application | Often zero copy |
| System calls | Needed to send and receive | Avoided |
| Timing | Variable, affected by scheduling | Consistent, low |

## Common technologies

| Technology | Description |
|---|---|
| Solarflare / AMD Onload and ef_vi | Accelerates standard socket applications (Onload) or offers a raw low level interface (ef_vi) |
| DPDK | Data Plane Development Kit, an open source framework for fast packet processing in user space |
| RDMA and InfiniBand | Remote direct memory access between machines |
| Exablaze / Cisco Nexus cards and similar | Ultra low latency network cards used in trading |
| AF_XDP | A Linux feature giving fast packet access with less complexity than full bypass |

Some, like Onload, can accelerate existing socket code with few changes; others require rewriting networking code.

## Busy polling

Kernel bypass applications usually spin in a loop, constantly checking for new packets, rather than sleeping until an interrupt wakes them. This uses a whole CPU core at 100% all the time but removes wake up delays. It pairs with pinning the thread to a dedicated core. See [CPU Affinity, NUMA and Cache Optimization](https://learn.tradelabsai.com/infrastructure/cpu-affinity/).

**Example: Where the microseconds go**
A firm measures the time from a market data packet arriving at its network card to the application seeing the decoded price. With the standard Linux network stack, the median is about 8 microseconds and the 99th percentile about 40 microseconds, because of interrupts and scheduling. With a kernel bypass library and a busy polling thread on an isolated core, the median falls to about 1.5 microseconds and the 99th percentile to about 3 microseconds. The improvement in the tail matters as much as the median, since the slowest moments are often the busiest and most valuable. These figures are illustrative; real results depend heavily on hardware and tuning. See [Exchange vs Receive Timestamps and Latency Measurement](https://learn.tradelabsai.com/programming/latency-measurement/).

## Trade offs

| Benefit | Cost |
|---|---|
| Much lower and steadier latency | Specialised network cards and licences |
| Fewer copies and system calls | Dedicated CPU cores burning at 100% |
| | More complex code and debugging |
| | Standard tools may not see the traffic |
| | Security and isolation handled by the application |

## Who uses it

Kernel bypass is common among market makers, high frequency firms and exchanges. It is overkill for nearly all retail and most institutional strategies, where network latency of milliseconds is perfectly acceptable. See [High-Frequency Trading](https://learn.tradelabsai.com/algo-trading/high-frequency-trading/).

## Beyond kernel bypass

When even microseconds are too slow, firms move logic into hardware itself, with FPGAs that parse market data and send orders without touching the CPU. See [FPGAs and Hardware Acceleration](https://learn.tradelabsai.com/infrastructure/fpgas-and-hardware-acceleration/).

## Frequently asked questions

### What is kernel bypass?

A technique that lets applications send and receive network packets directly with the network card, skipping the operating system's network stack to reduce latency.

### What is DPDK?

The Data Plane Development Kit, an open source set of libraries for fast packet processing in user space, widely used in networking and some trading systems.

### Does kernel bypass help retail traders?

No. Its gains are measured in microseconds, which matter only for latency critical professional strategies.

Next, learn how hardware acceleration goes even further in [FPGAs and Hardware Acceleration](https://learn.tradelabsai.com/infrastructure/fpgas-and-hardware-acceleration/).

## Continue learning

- Next lesson: [FPGAs and Hardware Acceleration](https://learn.tradelabsai.com/infrastructure/fpgas-and-hardware-acceleration/)
- Previous lesson: [Co-Location](https://learn.tradelabsai.com/infrastructure/co-location/)
- Related: [Co-Location](https://learn.tradelabsai.com/infrastructure/co-location/): Co location places trading servers inside or beside an exchange's data centre to cut latency. Learn how it works, what it costs, fairness rules and who needs it.
- Related: [Networking for Traders](https://learn.tradelabsai.com/infrastructure/networking-for-traders/): The networking basics behind trading: latency and distance, TCP versus UDP, multicast market data, packet loss, jitter and practical ways to improve connectivity.
- Related: [CPU Affinity, NUMA and Cache Optimization](https://learn.tradelabsai.com/infrastructure/cpu-affinity/): CPU affinity pins trading threads to specific cores so they are not interrupted or moved. Learn core isolation, NUMA, interrupts and how these cut latency spikes.
- Related: [FPGAs and Hardware Acceleration](https://learn.tradelabsai.com/infrastructure/fpgas-and-hardware-acceleration/): FPGAs run trading logic in custom hardware for nanosecond reaction times. Learn what FPGAs are, what firms put on them, the costs, the limits and alternatives.
- Related: [Exchange vs Receive Timestamps and Latency Measurement](https://learn.tradelabsai.com/programming/latency-measurement/): How to measure latency in a trading system: where to timestamp, tick to trade and order round trip, percentiles instead of averages and how to find bottlenecks.
