支持syscall/sysret的x86-64 Intel系统中,vanilla内核下最快的64位用户态系统调用是什么?
Great question! On vanilla x86-64 Intel Linux kernels using the syscall/sysret instruction pair, the "fastest" syscall that triggers a full user-kernel switch but does almost no additional work is an invalid syscall number—specifically one that gets rejected early in the kernel's syscall entry path without triggering heavyweight operations like signal delivery or error logging.
Why an invalid syscall number?
The kernel's syscall entry logic first checks if the provided syscall number falls within the valid range (i.e., less than NR_syscalls, which is typically between 400-500 on modern kernels). If it doesn't, the kernel immediately sets the return value to -EINVAL (error code 22) and jumps straight to the sysret path to return to user space. This skips all the logic that would execute a valid syscall: no function calls, no memory accesses to process state, no locks, no data structure manipulations—just the bare minimum to handle the switch and return an error.
What to avoid
You want to steer clear of invalid syscall numbers that might trigger slow paths:
- Negative syscall numbers can lead to different error handling in some kernel versions.
- Don't use numbers that map to deprecated or stub syscalls (some might still execute minimal logic).
Stick to a number well aboveNR_syscalls(like0x1000) for consistent fast rejection.
Comparison to "valid" lightweight syscalls
Even seemingly trivial syscalls like getpid() have more overhead: they require reading the pid field from the current task structure, which involves a memory access and potentially cache lookups. The invalid syscall skips all that—its only cost is the syscall/sysret switch plus a handful of CPU instructions in the kernel entry code.
Example code
Here's a quick C snippet using inline assembly to trigger this fast invalid syscall:
#include <stdio.h> int main() { long ret; // Syscall number 0x1000 is way above the valid range on most kernels asm volatile ( "syscall" : "=a"(ret) // Output: return value in rax : "a"(0x1000) // Input: invalid syscall number in rax : "rcx", "r11", "memory" // Clobbered registers per syscall ABI ); printf("Invalid syscall returned: %ld (-EINVAL = -22)\n", ret); return 0; }
This will output -22 as expected, and the entire syscall round-trip will be as fast as possible while still triggering a full user-kernel mode switch.
内容的提问来源于stack exchange,提问作者BeeOnRope

