如何通过内存分配或重定位更新页表的PMD与PTE表(基于Intel Xeon Gold 6142 4级页表架构)
Great question! Since you're working with an Intel Xeon Gold 6142 (which uses a 4-level page table: PML4 → PDPT → PMD → PTE), let's break down actionable solutions to trigger simultaneous PMD and PTE updates, and get your process holding those page table locks frequently.
First, a quick recap of x86-64 4-level page table boundaries (for 4KB small pages):
- Each PMD entry covers a 2MB virtual address range (bits 29–21 of the virtual address)
- Each PTE entry covers a 4KB page (bits 20–12)
- To hit both PMD and PTE layers, we need operations that modify PMD entries and their underlying PTE entries at the same time.
Solution 1: Cross-PMD Boundary mmap/munmap Cycles
This approach leverages memory allocations that span two adjacent PMD regions, forcing the kernel to create/destroy both PMD and PTE entries during mmap/munmap.
Step-by-Step Implementation:
Calculate a cross-PMD virtual address:
Pick an address that starts in the last 4KB of one PMD region and extends into the next. For example, if we target the PMD range0x00000000001FF000to0x00000000001FFFFF(last 4KB of a 2MB PMD block), we'll allocate2MB + 4KB(0x201000 bytes) to span into the next PMD (0x0000000000200000to0x00000000003FFFFF).Map the cross-PMD region:
UsemmapwithMAP_FIXEDto enforce the exact address, ensuring we span two PMDs:#include <sys/mman.h> #include <string.h> #include <stdio.h> #define PMD_SIZE (2 * 1024 * 1024) #define PAGE_SIZE (4 * 1024) #define ALLOC_SIZE (PMD_SIZE + PAGE_SIZE) int main() { // Target address: last 4KB of a PMD region void *addr = (void *)0x00000000001FF000; // Map cross-PMD anonymous memory void *map = mmap(addr, ALLOC_SIZE, PROT_READ | PROT_WRITE, MAP_PRIVATE | MAP_ANONYMOUS | MAP_FIXED, -1, 0); if (map == MAP_FAILED) { perror("mmap failed"); return 1; } // Touch all pages to trigger physical allocation (creates PTEs + PMDs) memset(map, 0, ALLOC_SIZE); // Unmap to destroy PMDs and PTEs munmap(map, ALLOC_SIZE); // Repeat this cycle to frequently hold PMD/PTE locks return 0; }Why this works:
- During
mmap+memset, the kernel must create a new PMD entry for the second 2MB region (since it was unused) and populate PTE entries for all 4KB pages in both PMD regions. - During
munmap, when the last PTE in a PMD region is destroyed, the kernel also tears down the PMD entry itself. This means the unmap operation modifies both PMD and PTE layers, requiring the kernel to hold their respective locks.
- During
Solution 2: Huge Page Split/Merge + Memory Migration
This method uses 2MB huge pages (which use PMD entries directly, no PTEs) and splits them into small pages, then migrates the small pages to trigger PTE updates, before merging back to huge pages to trigger PMD updates.
Step-by-Step Implementation:
Allocate a 2MB huge page:
void *huge_map = mmap(NULL, PMD_SIZE, PROT_READ | PROT_WRITE, MAP_PRIVATE | MAP_ANONYMOUS | MAP_HUGETLB | MAP_HUGE_2MB, -1, 0); if (huge_map == MAP_FAILED) { perror("huge page mmap failed"); return 1; } // Touch the huge page to confirm allocation memset(huge_map, 0, PMD_SIZE);Split the huge page into small pages:
Unmap a small portion of the huge page (e.g., the last 4KB) — this forces the kernel to split the huge page into 4KB small pages, creating a PTE page table under the PMD entry:munmap((char *)huge_map + PMD_SIZE - PAGE_SIZE, PAGE_SIZE);Migrate the small pages:
Use themove_pagessystem call to migrate the small pages to different physical addresses. This updates each PTE entry and requires holding PTE locks:#include <sys/mman.h> #include <sys/syscall.h> #include <unistd.h> // Helper to call move_pages long move_pages(int pid, unsigned long count, void **pages, const int *nodes, int *status, int flags) { return syscall(SYS_move_pages, pid, count, pages, nodes, status, flags); } // Prepare array of small page addresses void **pages = malloc((PMD_SIZE / PAGE_SIZE - 1) * sizeof(void *)); for (int i = 0; i < PMD_SIZE / PAGE_SIZE - 1; i++) { pages[i] = (char *)huge_map + i * PAGE_SIZE; } // Migrate pages to node 0 (adjust based on your NUMA setup) int nodes[PMD_SIZE / PAGE_SIZE - 1] = {0}; int status[PMD_SIZE / PAGE_SIZE - 1]; move_pages(getpid(), PMD_SIZE / PAGE_SIZE - 1, pages, nodes, status, 0);Merge back to huge pages:
Usemadviseto re-merge the small pages into a huge page, which destroys the PTE entries and updates the PMD entry:madvise(huge_map, PMD_SIZE - PAGE_SIZE, MADV_MERGEABLE); madvise(huge_map, PMD_SIZE - PAGE_SIZE, MADV_HUGEPAGE);Repeat the cycle:
Loop through split → migrate → merge to repeatedly trigger PMD and PTE updates, holding their locks.
Why this works:
- Split: Converts a PMD-only mapping to a PMD → PTE mapping, modifying both layers.
- Migration: Updates individual PTE entries as pages are moved to new physical addresses.
- Merge: Destroys the PTE table and restores the PMD direct mapping, modifying both layers again.
Key Notes
- Lock Behavior: Both solutions force the kernel to acquire the
page_table_lock(for PMD operations) and per-PTE spinlocks, exactly what you need to stress-test PMD/PTE lock contention. - Address Selection: For Solution 1, if
0x00000000001FF000is already in use, pick another address that aligns to the last 4KB of a free PMD region (use/proc/self/mapsto find free ranges). - Huge Page Prerequisites: For Solution 2, ensure your system has huge pages configured (check
/sys/kernel/mm/hugepages/hugepages-2048kB/nr_hugepages).
内容的提问来源于stack exchange,提问作者Mohammad Siavashi

