Linux内存区域标志:为何同时需要VM_WRITE与VM_MAYWRITE?
Great question—this is one of those Linux VM details that seems redundant at first glance, but it’s a deliberate design to handle permission flexibility and copy-on-write (COW) semantics cleanly. Let’s break this down step by step.
Core Definitions
First, let’s restate the flags clearly (from Understanding the Linux Virtual Memory Manager):
VM_WRITE: The memory region is currently writable (effective write permission).VM_MAYWRITE: The memory region is allowed to have VM_WRITE enabled (potential write permission).
Scenarios Where VM_MAYWRITE is Set but VM_WRITE is Not
There are several common cases where this combination occurs:
1. Post-Fork Copy-On-Write (COW) Protected Regions
When a process calls fork(), the kernel creates a near-identical copy of the parent’s address space. For private writable regions, instead of immediately copying all pages (which is expensive), the kernel:
- Marks the underlying page table entries as read-only for both parent and child.
- Keeps
VM_MAYWRITEset (since the region is inherently allowed to be writable). - Clears
VM_WRITEtemporarily to enforce the read-only state at the VM area level.
When either process attempts to write to these pages, a page fault occurs. The kernel checks VM_MAYWRITE: if it’s set, it performs COW (creates a private copy of the page for the writing process), then re-enables VM_WRITE and updates the page table to allow writes. If VM_MAYWRITE wasn’t set, the write would trigger a SIGSEGV (segmentation fault).
2. mprotect() Temporarily Revoking Write Permission
Suppose you have a memory region that was originally writable (both flags set). If you call mprotect(addr, len, PROT_READ), the kernel will:
- Clear
VM_WRITE(revoke current write access). - Keep
VM_MAYWRITEintact (since the region is still allowed to be made writable again later).
Later, if you call mprotect(addr, len, PROT_WRITE), the kernel will re-enable VM_WRITE because VM_MAYWRITE permits it. If VM_MAYWRITE had been cleared, this second mprotect() call would fail with EACCES.
3. Shared Memory with Restricted Per-Process Access
For shared memory regions (e.g., via shmget()), one process might choose to restrict its own write access while keeping the option to restore it. By clearing VM_WRITE but retaining VM_MAYWRITE, the process can safely read from the shared region, and later re-enable writes without needing to modify the shared memory’s global permissions.
How Kernel Behavior Differs in This State
When VM_MAYWRITE is set but VM_WRITE is not:
- Write attempts trigger a page fault: Instead of immediately crashing, the kernel checks
VM_MAYWRITE. If set, it may perform COW (for private regions) or re-enable write access (if permissions are adjusted). IfVM_MAYWRITEis unset, the write triggersSIGSEGV. - mprotect() can restore write access: You can call
mprotect()to setPROT_WRITEsuccessfully, asVM_MAYWRITEallows the transition. WithoutVM_MAYWRITE, this call would fail. - No implicit write access: The region behaves as read-only for normal operations, but the kernel remembers that write access is permitted if explicitly requested.
Why Not Use a Single Flag?
Combining these into one flag would break two critical Linux VM features:
- Flexible permission management: Separating "current access" from "allowed access" lets the kernel (and userspace) temporarily revoke permissions without permanently locking them down. This is useful for debugging, security policies, and dynamic access control.
- Efficient COW implementation: COW relies on knowing whether a region is capable of being writable, even if it’s currently read-only. A single flag couldn’t distinguish between "never writable" and "temporarily read-only for COW"—leading to incorrect page fault handling (e.g., crashing instead of performing COW).
- Consistency with Unix permission models: This split mirrors how file permissions work (e.g., a file might have
wpermission removed but still allow the owner to restore it). It’s a familiar pattern for permission layering.
内容的提问来源于stack exchange,提问作者Simple.guy

