Kernel Bypass and Low-Latency Networking
Kernel bypass lets trading software read network packets straight from the network card, skipping the operating system. Learn how it works and the trade offs.
When a network packet arrives at a normal server, the operating system kernel handles it first: interrupts fire, the packet is copied into kernel memory, passed through the network stack and finally copied again into the application. Each step costs time and adds unpredictable delays. Kernel bypass techniques let the trading application read packets directly from the network card's memory, skipping most of that path. For high frequency firms, kernel bypass can cut network processing from several microseconds to around a microsecond or less, with far more consistent timing.
The normal path versus kernel bypass#
| Step | Normal kernel networking | Kernel bypass |
|---|---|---|
| Packet arrives | Network card raises an interrupt | Application polls the card directly |
| Processing | Kernel network stack (IP, TCP or UDP) | User space library or the application |
| Copies | Into kernel buffers, then into the application | Often zero copy |
| System calls | Needed to send and receive | Avoided |
| Timing | Variable, affected by scheduling | Consistent, low |
Common technologies#
| Technology | Description |
|---|---|
| Solarflare / AMD Onload and ef_vi | Accelerates standard socket applications (Onload) or offers a raw low level interface (ef_vi) |
| DPDK | Data Plane Development Kit, an open source framework for fast packet processing in user space |
| RDMA and InfiniBand | Remote direct memory access between machines |
| Exablaze / Cisco Nexus cards and similar | Ultra low latency network cards used in trading |
| AF_XDP | A Linux feature giving fast packet access with less complexity than full bypass |
Some, like Onload, can accelerate existing socket code with few changes; others require rewriting networking code.
Busy polling#
Kernel bypass applications usually spin in a loop, constantly checking for new packets, rather than sleeping until an interrupt wakes them. This uses a whole CPU core at 100% all the time but removes wake up delays. It pairs with pinning the thread to a dedicated core. See CPU Affinity, NUMA and Cache Optimization.
Trade offs#
| Benefit | Cost |
|---|---|
| Much lower and steadier latency | Specialised network cards and licences |
| Fewer copies and system calls | Dedicated CPU cores burning at 100% |
| More complex code and debugging | |
| Standard tools may not see the traffic | |
| Security and isolation handled by the application |
Who uses it#
Kernel bypass is common among market makers, high frequency firms and exchanges. It is overkill for nearly all retail and most institutional strategies, where network latency of milliseconds is perfectly acceptable. See High-Frequency Trading.
Beyond kernel bypass#
When even microseconds are too slow, firms move logic into hardware itself, with FPGAs that parse market data and send orders without touching the CPU. See FPGAs and Hardware Acceleration.
Frequently asked questions#
What is kernel bypass?#
A technique that lets applications send and receive network packets directly with the network card, skipping the operating system's network stack to reduce latency.
What is DPDK?#
The Data Plane Development Kit, an open source set of libraries for fast packet processing in user space, widely used in networking and some trading systems.
Does kernel bypass help retail traders?#
No. Its gains are measured in microseconds, which matter only for latency critical professional strategies.
Next, learn how hardware acceleration goes even further in FPGAs and Hardware Acceleration.
3 quick questions on this lesson. Get them all right to finish it.
Turn on JavaScript to take the quiz.
Mentioned in
- Docker and KubernetesTrading Infrastructure
- Lock-Free Programming and Ring BuffersTrading Infrastructure
- Binary ProtocolsTrading Infrastructure
- Latency in TradingOrders and Execution
- High-Frequency TradingAlgorithmic Trading