execv后userfaultfd注册失败报ENOMEM,求解决方案
问题:父进程通过pidfd_getfd获取子进程的userfaultfd后,无法注册子进程execv后的内存
问题描述
我用Rust开发了一套程序,包含parent.rs和child.rs,执行流程如下:
- 父进程调用fork创建子进程
- 子进程创建
userfaultfd,并通过pidfd_getfd传递给父进程 - 子进程调用
execv运行child.rs - 子进程通过
mmap分配内存,把内存指针传递给父进程 - 父进程尝试用获取到的
userfaultfd句柄注册该内存,但最后一步失败,返回ENOMEM错误
我清楚虚拟内存机制下子进程的指针在父进程中无效,但认为该指针对已获取的userfaultfd句柄是有效的——因为创建该句柄时没设置O_CLOEXEC,而且已经通过/proc/{child_pid}/fd验证它处于打开状态。
当前临时方案及问题
我目前用LD_PRELOAD的hack方案临时解决,但存在不少弊端:
- 虽然解决了
cap_sys_ptrace能力与LD_PRELOAD的兼容问题,但实现非常繁琐 userfaultfd创建延迟,无法追踪LD_PRELOAD初始化前的内存分配- 方案不够可靠,存在安全隐患
我的实际应用中用ptrace可以灵活控制子进程,希望找到更合理的方法,让userfaultfd能正常注册子进程的内存。
相关代码及运行命令
Cargo.toml
[package] name = "..." version = "0.1.0" edition = "2021" [dependencies] nix = "0.26.2" pidfd = "0.2.4" pidfd_getfd = { version = "0.2.1", features = ["nightly"] } pipe-channel = "1.3.0" rustix = { version = "0.37.3", features = ["mm"] } userfaultfd = { version = "0.5.1", features = ["linux4_14", "linux5_7"] }
examples/parent.rs
use { nix::unistd, rustix::fd::{AsRawFd, FromRawFd}, std::{ ffi::{self, CString}, mem, }, userfaultfd::Uffd, }; fn main() { let child_name = std::env::args().nth(1).expect("Expected argument"); // Fork and execute the child let (mut uffd_tx, mut uffd_rx) = pipe_channel::channel(); let (mut ready_tx, mut ready_rx) = pipe_channel::channel(); let child_pid = match unsafe { unistd::fork() }.expect("Unable to fork") { unistd::ForkResult::Parent { child } => child, unistd::ForkResult::Child => { // Open the uffd and send it to the parent // Note: We forget it so it doesn't get closed. let uffd = userfaultfd::UffdBuilder::new() .close_on_exec(false) .user_mode_only(true) .non_blocking(false) .create() .expect("Unable to create uffd"); uffd_tx.send(uffd.as_raw_fd()).expect("Unable to send uffd"); mem::forget(uffd); // Wait until the monitor process is ready ready_rx.recv().expect("Unable to wait for parent"); // Then execute the child println!("Executing child"); let path = CString::new(child_name.clone()).unwrap(); let args = [CString::new(child_name).unwrap()]; unistd::execv(&path, &args).expect("Unable to `execv`"); unreachable!(); }, }; // Open a pid_fd for the child process let child_pidfd = unsafe { pidfd::PidFd::open(child_pid.as_raw(), 0) }.expect("Unable to allocate pidfd for child process"); // Receive the uffd from the child let child_uffd_fd = uffd_rx.recv().expect("Unable to receive uffd"); let uffd_fd = unsafe { pidfd_getfd::pidfd_getfd(child_pidfd.as_raw_fd(), child_uffd_fd, 0) }; let uffd = unsafe { Uffd::from_raw_fd(uffd_fd) }; // Tell the child we're ready to execute ready_tx.send(()).expect("Unable to send parent event"); // Then read the pointer it wrote std::thread::sleep(std::time::Duration::from_secs(1)); let page = std::fs::read("ptr").expect("Unable to read pointer"); let page = page.try_into().expect("File wasn't the right size"); let page = usize::from_le_bytes(page); let page = page as *mut ffi::c_void; // Prove the pointer is on the process's maps let memory_map = std::fs::read_to_string(format!("/proc/{}/maps", child_pid.as_raw())).expect("Unable to read memory maps"); assert!(memory_map.contains(&format!("{:x}", page as usize))); // Then try to register uffd.register(page, 4096).expect("Unable to register dummy pointer"); }
examples/child.rs
use {rustix::mm, std::ptr}; pub fn main() { // Allocate the page println!("Child: Allocating"); let page = unsafe { mm::mmap_anonymous( ptr::null_mut(), 4096, mm::ProtFlags::READ | mm::ProtFlags::WRITE, mm::MapFlags::PRIVATE, ) .expect("Unable to allocate page") }; // Write to file println!("Child: Writing to file"); std::fs::write("ptr", (page as usize).to_le_bytes()).expect("Unable to write"); println!("Child: Sleeping"); loop { std::thread::park(); } }
运行命令
cargo build --examples && ./target/debug/examples/parent ./target/debug/examples/child
解决方案
核心问题分析
你遇到的ENOMEM错误根源在于:userfaultfd是和进程地址空间绑定的。子进程在fork后创建的userfaultfd,关联的是fork后的子进程地址空间;但当子进程调用execv后,原地址空间被完全替换成新程序的地址空间,此时原来的userfaultfd已经无法关联到新的地址空间了——即便你通过pidfd_getfd把句柄传到父进程,这个句柄对应的userfaultfd依然绑定的是execv前的旧地址空间,无法操作新地址空间的内存。
正确实现思路
既然你已经在使用ptrace控制子进程,最合理的方案是在子进程execv完成后,通过ptrace注入系统调用,在新的地址空间中创建userfaultfd,再传递给父进程。具体步骤如下:
- 父进程
fork子进程后,立即用ptrace跟踪子进程(设置PTRACE_SEIZE或PTRACE_TRACEME)。 - 子进程调用
execv后,会触发SIGTRAP信号,父进程捕获到这个信号后,确认子进程已进入新程序的地址空间。 - 父进程通过
ptrace向子进程注入userfaultfd系统调用,创建新的userfaultfd句柄。 - 父进程通过
pidfd_getfd获取这个新的userfaultfd句柄。 - 后续父进程就可以用这个句柄注册子进程
execv后分配的内存了。
关键代码调整示例(基于现有代码)
修改parent.rs的核心逻辑
// 替换原fork后的逻辑,加入ptrace跟踪 let child_pid = match unsafe { unistd::fork() }.expect("Unable to fork") { unistd::ForkResult::Parent { child } => child, unistd::ForkResult::Child => { // 子进程开启ptrace跟踪自身 unsafe { nix::sys::ptrace::ptrace(nix::sys::ptrace::Request::PTRACE_TRACEME, nix::unistd::Pid::from_raw(0), std::ptr::null(), std::ptr::null()) }.expect("Failed to set PTRACE_TRACEME"); // 发送SIGSTOP让父进程捕获,确保跟踪生效 unsafe { nix::sys::signal::kill(nix::unistd::getpid(), nix::sys::signal::Signal::SIGSTOP) }.expect("Failed to send SIGSTOP"); // 执行child程序 println!("Executing child"); let path = CString::new(child_name.clone()).unwrap(); let args = [CString::new(child_name).unwrap()]; unistd::execv(&path, &args).expect("Unable to `execv`"); unreachable!(); }, }; // 父进程等待子进程的SIGSTOP信号 unsafe { nix::sys::wait::waitpid(child_pid, None, nix::sys::wait::WaitPidFlag::WUNTRACED) }.expect("Failed to wait for child"); // 恢复子进程执行,等待execv后的SIGTRAP(execv完成会触发) unsafe { nix::sys::ptrace::ptrace(nix::sys::ptrace::Request::PTRACE_CONT, child_pid, std::ptr::null(), std::ptr::null()) }.expect("Failed to continue child"); let wait_status = unsafe { nix::sys::wait::waitpid(child_pid, None, nix::sys::wait::WaitPidFlag::WUNTRACED) }.expect("Failed to wait for exec trap"); assert!(matches!(wait_status, nix::sys::wait::WaitStatus::Stopped(_, nix::sys::signal::Signal::SIGTRAP))); // 此时子进程已进入新地址空间,注入userfaultfd系统调用(x86_64架构示例) // 获取子进程寄存器状态 let mut regs = unsafe { nix::sys::ptrace::ptrace(nix::sys::ptrace::Request::PTRACE_GETREGS, child_pid, std::ptr::null(), std::ptr::null()) }.expect("Failed to get registers"); // 设置rax为x86_64的userfaultfd系统调用号(383),rdi为创建参数 regs.rax = 383; regs.rdi = (userfaultfd::UffdFlags::USER_MODE_ONLY).bits() as u64; // 写入寄存器并触发系统调用 unsafe { nix::sys::ptrace::ptrace(nix::sys::ptrace::Request::PTRACE_SETREGS, child_pid, std::ptr::null(), ®s) }.expect("Failed to set registers"); unsafe { nix::sys::ptrace::ptrace(nix::sys::ptrace::Request::PTRACE_SYSCALL, child_pid, std::ptr::null(), std::ptr::null()) }.expect("Failed to trigger syscall"); // 等待系统调用完成 let wait_status = unsafe { nix::sys::wait::waitpid(child_pid, None, nix::sys::wait::WaitPidFlag::WUNTRACED) }.expect("Failed to wait for syscall completion"); assert!(matches!(wait_status, nix::sys::wait::WaitStatus::Stopped(_, nix::sys::signal::Signal::SIGTRAP))); // 获取系统调用返回的uffd句柄 let regs = unsafe { nix::sys::ptrace::ptrace(nix::sys::ptrace::Request::PTRACE_GETREGS, child_pid, std::ptr::null(), std::ptr::null()) }.expect("Failed to get registers"); let child_uffd_fd = regs.rax as i32; // 用pidfd_getfd获取这个新的uffd let child_pidfd = unsafe { pidfd::PidFd::open(child_pid.as_raw(), 0) }.expect("Unable to allocate pidfd for child process"); let uffd_fd = unsafe { pidfd_getfd::pidfd_getfd(child_pidfd.as_raw_fd(), child_uffd_fd, 0) }; let uffd = unsafe { Uffd::from_raw_fd(uffd_fd) }; // 恢复子进程正常执行 unsafe { nix::sys::ptrace::ptrace(nix::sys::ptrace::Request::PTRACE_CONT, child_pid, std::ptr::null(), std::ptr::null()) }.expect("Failed to continue child"); // 后续读取指针并注册的逻辑保持不变 std::thread::sleep(std::time::Duration::from_secs(1)); let page = std::fs::read("ptr").expect("Unable to read pointer"); // ... 原注册逻辑 ...
注意事项
- 系统调用号需要根据目标架构调整(x86_64是383,arm64是277),可通过
man 2 userfaultfd查看。 - 注入系统调用时要严格遵循对应架构的参数传递规则(x86_64用rdi/rsi等寄存器,arm64用x0/x1等)。
- 必须等待子进程
execv完成并触发SIGTRAP后,再创建userfaultfd,确保关联的是新地址空间。
内容的提问来源于stack exchange,提问作者Filipe Rodrigues
相关产品推荐
相关产品推荐

