关于CPU仅识别虚拟地址时页错误处理器填充页表项物理地址的技术疑问
Awesome questions—this stuff is the nitty-gritty of virtual memory, and it’s easy to get confused when you’re first wrapping your head around how MMUs and page faults interact. Let’s break this down one piece at a time:
First, the Critical Premise: The Kernel Can Access Physical Memory Directly
Before diving into your specific questions, let’s establish the foundation: the operating system’s kernel runs in a privileged mode (like Ring 0 on x86, EL1 on ARM) where it’s not bound by the same virtual address rules as user-space programs. It can either:
- Use a direct 1:1 mapped virtual address space (most modern OSes reserve a chunk of virtual memory to mirror physical memory exactly—e.g., the top 1GB on x86 maps directly to the first 1GB of physical RAM), or
- Temporarily disable the MMU entirely (on some architectures) to access physical addresses directly.
Without this ability, the kernel couldn’t manage memory at all—so this is non-negotiable for how page fault handling works.
1. How does the page fault handler fill page table entries with physical addresses if the CPU only understands virtual addresses?
Since the page fault handler is part of the kernel, it doesn’t rely on the user-space page tables that caused the fault. Here’s how it works:
- The kernel uses its direct-mapped virtual address space to access the page table itself (which lives in physical memory). For example, if the page table is at physical address
0x100000, the kernel can access it via a corresponding virtual address in its direct-mapped region (like0xffff800000100000on x86). - It writes the physical page’s address directly into the correct page table entry, along with necessary flags (e.g., "present", "user-accessible", "read-write").
- Finally, it tells the MMU to refresh its TLB (Translation Lookaside Buffer)—the cache of recent address mappings—so the new entry is used for future accesses.
No need for the CPU to "understand" physical addresses here; the kernel’s virtual address space acts as a direct window to physical memory.
2. Where does the page fault handler get physical addresses in the first place?
The kernel owns and manages all physical memory via a physical page frame allocator (like the buddy system in Linux, or the zone allocator in Windows). This allocator keeps track of every 4K (or larger) physical page in the system:
- It maintains metadata (stored in physical memory, accessible via the kernel’s direct mapping) that tracks which pages are free, which are in use, and their exact physical addresses.
- When the handler needs a physical page, it simply asks the allocator for a free frame. The allocator checks its free lists, grabs an available page, and returns its physical address (e.g.,
0x40000000).
For pages that need to be loaded from disk, the kernel also keeps metadata (like page table entries with disk addresses, or file inode data) that tells it exactly where the virtual page is stored on disk—so it knows which blocks to read into the new physical page.
3. Does the CPU use physical addresses directly during page fault handling?
Sort of—but not in the way you might expect. Most modern kernels stick to using their direct-mapped virtual address space, so the CPU is still executing instructions that reference virtual addresses. But these virtual addresses map 1:1 to physical addresses, so accessing them is functionally identical to using physical addresses directly.
Some architectures do allow the kernel to disable the MMU temporarily, which lets the CPU use physical addresses for memory operations. But this is rare in modern systems because it adds overhead (you have to re-enable the MMU and flush caches afterward) and the direct-mapped region is sufficient for most tasks.
The key takeaway: The CPU doesn’t suddenly switch to a "physical address mode"—it’s still using virtual addresses, but the kernel’s virtual addresses give it unfiltered access to physical memory without relying on the broken user-space page tables that triggered the fault.
4. How does the handler get a 4K physical page and its address when loading from disk?
Let’s walk through this scenario step by step:
- Grab a free physical page: The kernel calls its page allocator (e.g.,
__get_free_page()in Linux) to request an unused 4K frame. If memory is tight, the allocator might first evict a less-used page (using a replacement algorithm like LRU) to free up space—writing that page back to disk if it’s been modified. - Read disk data into the physical page: The kernel uses the disk controller’s interface (like NVMe or SCSI) to read the corresponding 4K block from disk into the new physical page. It knows exactly where the virtual page lives on disk (from metadata like the page table’s "disk address" field) and sends the correct LBA (Logical Block Address) to the disk controller.
- Update the page table: Using its direct-mapped virtual address, the kernel locates the page table entry for the faulting virtual address. It writes the physical page’s address into the entry, sets the "present" flag and other permissions, and marks the entry as valid.
- Refresh the TLB: The kernel tells the MMU to invalidate the old (invalid) entry for the virtual address in the TLB, so the next access uses the new mapping.
- Resume execution: The CPU re-runs the instruction that caused the page fault. This time, the MMU finds the valid page table entry, translates the virtual address to the physical address, and the access succeeds.
内容的提问来源于stack exchange,提问作者curious-carp

