TradeLabs AILearn

Kernel Bypass and Low-Latency Networking

Kernel bypass lets trading software read network packets straight from the network card, skipping the operating system. Learn how it works and the trade offs.

Advanced3 min readUpdated 3 Oct 2026
Markdown
Read firstCo-Location
Lesson 11 of 16

When a network packet arrives at a normal server, the operating system kernel handles it first: interrupts fire, the packet is copied into kernel memory, passed through the network stack and finally copied again into the application. Each step costs time and adds unpredictable delays. Kernel bypass techniques let the trading application read packets directly from the network card's memory, skipping most of that path. For high frequency firms, kernel bypass can cut network processing from several microseconds to around a microsecond or less, with far more consistent timing.

The normal path versus kernel bypass#

StepNormal kernel networkingKernel bypass
Packet arrivesNetwork card raises an interruptApplication polls the card directly
ProcessingKernel network stack (IP, TCP or UDP)User space library or the application
CopiesInto kernel buffers, then into the applicationOften zero copy
System callsNeeded to send and receiveAvoided
TimingVariable, affected by schedulingConsistent, low

Common technologies#

TechnologyDescription
Solarflare / AMD Onload and ef_viAccelerates standard socket applications (Onload) or offers a raw low level interface (ef_vi)
DPDKData Plane Development Kit, an open source framework for fast packet processing in user space
RDMA and InfiniBandRemote direct memory access between machines
Exablaze / Cisco Nexus cards and similarUltra low latency network cards used in trading
AF_XDPA Linux feature giving fast packet access with less complexity than full bypass

Some, like Onload, can accelerate existing socket code with few changes; others require rewriting networking code.

Busy polling#

Kernel bypass applications usually spin in a loop, constantly checking for new packets, rather than sleeping until an interrupt wakes them. This uses a whole CPU core at 100% all the time but removes wake up delays. It pairs with pinning the thread to a dedicated core. See CPU Affinity, NUMA and Cache Optimization.

Trade offs#

BenefitCost
Much lower and steadier latencySpecialised network cards and licences
Fewer copies and system callsDedicated CPU cores burning at 100%
More complex code and debugging
Standard tools may not see the traffic
Security and isolation handled by the application

Who uses it#

Kernel bypass is common among market makers, high frequency firms and exchanges. It is overkill for nearly all retail and most institutional strategies, where network latency of milliseconds is perfectly acceptable. See High-Frequency Trading.

Beyond kernel bypass#

When even microseconds are too slow, firms move logic into hardware itself, with FPGAs that parse market data and send orders without touching the CPU. See FPGAs and Hardware Acceleration.

Frequently asked questions#

What is kernel bypass?#

A technique that lets applications send and receive network packets directly with the network card, skipping the operating system's network stack to reduce latency.

What is DPDK?#

The Data Plane Development Kit, an open source set of libraries for fast packet processing in user space, widely used in networking and some trading systems.

Does kernel bypass help retail traders?#

No. Its gains are measured in microseconds, which matter only for latency critical professional strategies.

Next, learn how hardware acceleration goes even further in FPGAs and Hardware Acceleration.

Check your understanding

3 quick questions on this lesson. Get them all right to finish it.

Turn on JavaScript to take the quiz.

Finished this lesson?Sign in to save your progress across devices.
Next lessonFPGAs and Hardware AccelerationFPGAs run trading logic in custom hardware for nanosecond reaction times. Learn what FPGAs are, what firms put on them, the costs, the limits and alternatives.

Mentioned in