基于QEMU+GDB追踪用户态Hello World至内核态调用流程
Great question—let's walk through exactly how to trace your simple C program all the way from user-space execution into the kernel's system call handling, using your existing QEMU+GDB setup.
Step 1: Write & Compile the Test Program
First, create a minimal hello.c with the required printf call, and compile it with debug symbols so GDB can properly trace user-space code (we'll use static linking to avoid dynamic library complexities):
#include <stdio.h> int main() { printf("hello world\n"); return 0; }
Compile it with:
gcc -g -static hello.c -o hello
The -g flag adds debug symbols, and -static bundles all dependencies into the executable, making the call chain far easier to follow without chasing dynamic library calls.
Step 2: Launch QEMU for Debugging
You likely already know this, but let's align on the QEMU command to ensure debugging works smoothly:
qemu-system-x86_64 -kernel ./vmlinuz -initrd ./initramfs.img -s -S -append "console=ttyS0 nokaslr"
-sspins up a GDB server on port 1234-Spauses the VM until GDB connects (so we don't miss the start of our program)nokaslrdisables kernel address space randomization, keeping breakpoint addresses stable and predictable
Step 3: Connect GDB & Load Symbols
Open a new terminal, launch GDB, load your kernel symbols, connect to QEMU, and prepare to load your user program's symbols:
gdb ./vmlinux target remote :1234
Now resume the VM with c, then switch to your QEMU terminal and run ./hello. Hit Ctrl+C in GDB to pause execution, then load the user program's symbols using its base address:
# First, find the base address of hello in QEMU: run `cat /proc/<PID>/maps` where <PID> is hello's process ID add-symbol-file ./hello 0x<HELLO_BASE_ADDRESS>
This tells GDB how to map user-space memory addresses to your program's debug symbols.
Step 4: Trace the Full Call Chain
Let's start tracing from the very beginning of main:
- Set a breakpoint at the user-space
mainfunction:b main - Resume execution with
c—the program will stop at the first line ofmain. - Use
stepi(step one instruction) to walk through each call, orbt(backtrace) at any point to see the current call stack.
User-Space Call Sequence
You'll first traverse this user-space chain:
main()→ callsprintf("hello world\n")printf()→ delegates tovfprintf()(the core formatted print implementation)vfprintf()→ eventually calls the user-spacewrite()wrapper function, which prepares to trigger the system call
The Critical User-to-Kernel Switch
When the user-space write() executes the syscall instruction (x86_64) or int 0x80 (legacy x86), GDB will show a clear mode switch (look for output like [Switching to Thread 1.1] or a change to Kernel mode). This is where the CPU switches from unprivileged user mode to privileged kernel mode.
Kernel-Space Call Sequence
From here, the kernel handles the system call with this typical chain (x86_64 example):
entry_SYSCALL_64(kernel system call entry stub) → saves user-space registers and contextdo_syscall_64→ looks up the system call table using the syscall number (forwrite, that's__NR_write) to findsys_writesys_write→ the kernel's top-level write handlervfs_write→ the virtual filesystem (VFS) layer handler, which abstracts filesystem specifics__vfs_write→ dispatches to the actual filesystem-specific write function (e.g.,ext4_file_write_iterif you're using an ext4 filesystem)- Once the write completes, the kernel restores user-space context and executes a
sysretinstruction to return to user-space
After returning to user space, execution continues up the call chain: user-space write() → vfprintf() → printf() → main(), which exits normally.
Bonus: Tracing Page Faults (Leveraging Your Existing __do_page_fault Breakpoint)
Since you're already familiar with setting breakpoints at __do_page_fault, you can trace a page fault triggered by your program—for example, when the kernel first loads the printf code into physical memory. Set your breakpoint:
b __do_page_fault
When the program hits this breakpoint, run bt to see the full stack: it will show the user-space instruction that triggered the fault, followed by the kernel's page fault handling chain (__do_page_fault → handle_mm_fault → __handle_mm_fault → etc.).
Key Takeaways
- The user-to-kernel transition relies on a CPU-specific privileged instruction that triggers a mode switch, not a regular function call.
- System calls are routed through the kernel's system call table, which maps user-space syscall numbers to kernel handler functions.
- Debugging both user and kernel space requires loading symbols for both the user program and the kernel in GDB—this lets you seamlessly trace across the mode boundary.
内容的提问来源于stack exchange,提问作者Debashish

