Direct Kernel Invocation Overview
- Direct Kernel Invocation is a method to call OS kernel functions directly, bypassing conventional system calls to achieve lower latency and enhanced throughput.
- It uses techniques like unikernel hybrid linking, in-kernel eBPF execution, and cache-coherent remote invocation to optimize system responsiveness and resource utilization.
- The approach involves trade-offs in safety and flexibility, requiring secure linking, process isolation, and careful memory management to mitigate potential risks.
Direct Kernel Invocation refers to architectural, programming, or system-level mechanisms by which code in one context (user level, user agent, external device, or remote peer) synchronously calls into an operating system kernel or kernel-adjacent execution context, eliminating conventional privilege boundary crossings, entry/exit scaffolding, or multi-stage control flow. It encompasses innovations such as statically-linked kernel-userspace hybrid binaries, eBPF and remote/heterogeneous direct invocation, cache-coherent device protocols, and code-mobility APIs for injected execution on smart devices or remote endpoints. These enable latency reduction, tail-latency elimination, and increase IOPS for workloads ranging from in-kernel storage chains to serverless RPC, dataflow offload, smart storage, and in-band accelerator code deployment.
1. Kernel/User Boundary Elimination via Hybrid Linking
Direct kernel invocation at the binary and ABI level is exemplified by Unikernel Linux (UKL), which enables a user-level process and the Linux kernel to be compiled, linked, and executed as a single hybrid binary. UKL reconfigures Linux via a Kconfig target (CONFIG_UKL), modifies the glibc build to generate syscall wrappers that issue direct calls rather than syscalls/sysrets, disables the red zone, forces kernel model compilation, and adjusts linker scripts to merge user and kernel symbols with relocation (prefixing user symbols with "ukl_") (Raza et al., 2022).
Privilege and context tracking are managed using a per-task flag (ukl_mode) in task_struct, allowing code running at supervisor privilege but logically in "user" context to access kernel APIs either through conventional control paths (with RCU, scheduler checks) or optionally bypassing those (UKL_BYP, UKL_RET, etc.) for lower latency. In the "deep shortcut" case, application code can directly call internal kernel routines (e.g., tcp_sendmsg_internal) bypassing all syscall machinery.
Classic syscall transitions () are replaced with a direct function call model (), empirically yielding reductions such as , , and . Table 1 below summarizes end-to-end improvements in a Redis server benchmark (Raza et al., 2022):
| Configuration | 99% Lat (ms) | Throughput (op/s) | Δ Thr vs Linux |
|---|---|---|---|
| Linux 5.14 | 3.26 | 6,375,200 | — |
| UKL base | 3.25 | 6,479,200 | +1.6% |
| UKL_RET_BYP (bypass+ret) | 2.91 | 7,154,680 | +12% |
| UKL_RET_BYP+TCP-shortcut | 2.54 | 8,022,540 | +26% |
These demonstrate tail-latency and throughput gains, attributing the improvement to direct invocation and bypassing of syscall machinery (Raza et al., 2022).
2. In-Kernel Direct Invocation via eBPF and Storage Paths
Direct kernel invocation can also be implemented through programmable in-kernel code, specifically eBPF, to eliminate user/kernel crossings for dependent operations. This is illustrated in BPF-enabled storage stacks, where user-supplied eBPF programs are loaded into the kernel (via bpf(BPF_PROG_LOAD,…)) and attached to specific file descriptors through a custom ioctl (STORAGE_BPF_ATTACH). The kernel then invokes the eBPF code directly at two critical hook points:
- Syscall-dispatch hook: Placed after read()/io_uring entry, eliminates round-trips for chained operations.
- NVMe-driver completion hook: Enables the driver IRQ handler to reissue follow-on I/Os (e.g., B-tree lookups) without traversing additional kernel or file system layers (Wu et al., 2021).
The user/kernel boundary elimination is quantitatively significant: the NVMe-driver hook yields up to 2.5× IOPS over baseline and 49% reduction in median latency for dependent B-tree lookups.
Layer-wise breakdown in 512 B random read:
| Layer | Latency | % of total |
|---|---|---|
| User↔Kernel crossing | 351 ns | 5.6% |
| read() syscall | 199 ns | 3.2% |
| ext4 | 2006 ns | 32% |
| bio layer | 379 ns | 6.0% |
| NVMe driver | 113 ns | 1.8% |
| Storage device | 3224 ns | 51.4% |
The exokernel-inspired policy attaches per-file block extent caching for safe, offset-scoped block reads, and adds a per-process “chain counter” to limit the length of in-kernel traversals (Wu et al., 2021).
3. RPC-Style Direct Invocation Using Cache Coherent Interconnects
Direct kernel invocation can be generalized to RPC-style offload to accelerators across cache-coherent interconnects. On the Enzian platform, an open, programmable cache-coherence protocol (ECI) enables the FPGA to participate as a coherence home node, observing all cache line exchanges and intervention requests. A direct invocation comprises:
- CPU writes request arguments to a BAR-mapped cache line (B), issues a DMB barrier, and reads from line A.
- The FPGA “directory controller” observes the cache transaction, fetches inputs, evaluates the kernel logic, writes outputs, and responds so the CPU load unblocks with results (Ruzhanskaia et al., 2024).
Latency models for N-byte payloads, . Empirically, 16 B to 128 B RPCs exhibit ≈900 ns median with programmed I/O over ECI, substantially outperforming DMA over PCIe (minimum 2,350 ns). Tail-latency is virtually eliminated compared to interrupt or queue-driven models.
| Payload | DMA-PCIe | PIO-PCIe | PIO-ECI |
|---|---|---|---|
| 16 B | 2350 ns | 900 ns | 900 ns |
| 1 KiB | 2400 ns | 4600 ns | 1500 ns |
For receive paths (NIC use-case): RX median with PIO-ECI is 1.05 μs vs. 65.4 μs for DMA-PCIe, with complete elimination of high-percentile tail latency (Ruzhanskaia et al., 2024).
4. Remote and Heterogeneous Direct Invocation via Code Injection
Distributed and heterogeneous system designs increasingly support direct remote kernel invocation via code injection APIs. The UCX framework introduces an "ifunc" model whereby user code (as a dynamic library) and its arguments are packaged together and moved over RDMA to a remote endpoint (DPU, CSD, SmartNIC, or remote CPU). Invocation proceeds as follows (Peña et al., 2021):
- UCX registers the ifunc, dynamically links required symbols, and creates a message frame containing code and data.
- The message is sent by zero-copy RDMA PUT to a pre-mapped, RWX buffer region at the target.
- The remote side polls for an incoming message, validates, relinks GOT pointers, and directly executes the injected main() callback in local address space.
Latency and throughput crossovers are observed as payload grows (for small payloads, classic Active Messages outperform; for large data/code, ifunc closes the gap or surpasses), with a notable decrease in round-trips due to one-sided, immediate execution.
| Payload | UCX AM (μs) | ifunc (μs) | Δ | UCX AM (Mmsg/s) | ifunc (Mmsg/s) | Speedup |
|---|---|---|---|---|---|---|
| 1 KB | 3.5 | 4.0 | +14% | 4.2 | 2.3 | 0.55× |
| 16 KB | 9.8 | 9.4 | –4% | n/a | n/a | n/a |
| 1 MB | 80.2 | 52.0 | –35% | 0.012 | 0.022 | 1.8× |
This framework breaks the SPMD and static handler ID limitations of traditional HPC and RPC systems, supporting on-the-fly deployment of new kernels, late binding, dynamic code placement, and device heterogeneity (Peña et al., 2021).
5. Mechanisms, Trade-Offs, and Safety Considerations
Direct kernel invocation—whether through hybrid binaries, in-kernel interpreters, or remote code mobility—trades off safety and flexibility for performance:
- UKL mandates static linkage and supervisor privilege; thus, co-running untrusted processes require hypervisor mediation. Dynamic linking, dlopen support, and broader architecture support remain research frontiers.
- eBPF/exokernel approaches must manage file extent cache invalidation and restrict block address access to ensure safety and fairness (per-process chain limits), borrowing exokernel primitives for safe mapping and capability scoping (Wu et al., 2021).
- Cache-coherent accelerator protocols demand robust directory controller logic, fairness in conflicting line requests, and careful management of cache/L1/L2/TAD resources.
- Remote code injection APIs rely on explicit buffer management, securely loading and verifying code, and enforcing execution within controlled, dynamically allocated contexts; one-sided RDMA also raises isolation challenges.
Potential future optimizations common to all approaches include widespread link-time optimization across user and kernel code, zero-copy I/O, fine-grained control over critical path synchronization, and architectural extensions for safe direct invocation primitives.
6. Impact and Future Research Directions
Direct kernel invocation brings systematic tail-latency reduction, deterministic microsecond-level I/O invocation (e.g., ≤20 ns call overhead in UKL), and enables new models of kernel and accelerator composition in both general-purpose and data-centric environments. It supports incremental transformation from monolithic OSes toward exokernel, unikernel, and distributed microkernel paradigms without abandoning legacy ecosystems (Raza et al., 2022, Wu et al., 2021, Ruzhanskaia et al., 2024, Peña et al., 2021).
Current research focuses on:
- Wider architectural coverage (non-x86 targets)
- Injection and safe dynamic linking in kernel and device contexts
- Further fine-tuning of direct/injected control paths
- Security mechanisms for isolation in mixed-trust environments
This suggests direct kernel invocation, in its multiple emergent forms, will remain central to next-generation OS, storage, and distributed runtime system architectures, providing an essential substrate for high-throughput, low-latency, and safely extensible compute paths across CPU, device, and network boundaries.