关于用户程序触发系统调用后从ISR到系统调用表的执行流程问询
Great question—this is one of those kernel internals details that’s easy to skim over but really unlocks how userland programs talk to the kernel. Let’s break this down clearly, using x86 Linux as a concrete example since it’s the most widely used architecture for such discussions.
First, a quick clarification to clear up a common misconception: the ISR for system calls is not a separate "generic interrupt handler" that finishes execution before touching the syscall table. Instead, this ISR is the dedicated bridge between userland and the syscall table—its entire core purpose is to handle the handoff to the correct syscall implementation.
Pre-Requisite: How Userland Triggers a Syscall
- A user program triggers a software trap (trap) via instructions like
int 0x80(legacy x86),syscall(modern x86_64), orsysenter. - The CPU automatically switches to kernel mode, saves critical user-state registers (like the syscall number stored in
%eax/%rax) to the kernel stack, and jumps to the entry point defined in the Interrupt Descriptor Table (IDT).
The ISR's Core Flow (The "Missing" Steps You're Asking About)
Once the CPU jumps to the system call ISR (e.g., syscall_entry on x86_64), here’s exactly what happens to get to the syscall table:
- Save full user context: Beyond the registers the CPU saves automatically, the ISR pushes all remaining user-state registers (like
%rbx,%rbp,%r12-%r15) to the kernel stack. This ensures we can fully restore userland state after the syscall completes. - Validate the syscall number:
- The ISR reads the syscall number from the pre-saved register (e.g.,
%raxon x86_64). - It checks if this number falls within the valid range (i.e., less than
NR_syscalls, the total number of registered syscalls). If not, it jumps to an error handler that returns-ENOSYS(invalid syscall) to userland.
- The ISR reads the syscall number from the pre-saved register (e.g.,
- Jump to the syscall table entry:
- The syscall table is a global array of function pointers (
sys_call_tablein Linux), where each index maps directly to a syscall implementation (e.g., index 1 =sys_exit, index 3 =sys_read). - The ISR uses an indirect jump instruction to call the function at the index matching the syscall number. For x86_64, this looks like:
call *sys_call_table(,%rax,8)—the8accounts for each pointer being 8 bytes wide.
- The syscall table is a global array of function pointers (
- Execute the syscall logic: At this point, control is fully transferred to the specific syscall implementation (e.g.,
sys_read), which performs the actual kernel-side work (like accessing the file system or hardware).
Bonus: What Happens After the Syscall Finishes?
Once the syscall function returns, its return value is stored in a designated register (again, %rax on x86_64). The ISR then:
- Restores all saved user-state registers from the kernel stack.
- Uses a mode-switching instruction (like
sysretfor x86_64 oriretfor legacy x86) to switch back to user mode and return control to the original user program.
Why You Might Have Confused the Flow
If you’ve worked with hardware interrupts before, their ISRs do finish most of their work before returning—but system call ISRs are unique. They’re not handling a hardware event; they’re acting as a router, taking the user’s request and directing it to the correct kernel function via the syscall table.
内容的提问来源于stack exchange,提问作者Fabio

