锁前缀指令能否为弱序内存访问提供屏障?x86架构技术问询
Great question—let’s dig into the Intel Software Developer Manual (SDM) to get a definitive answer, since that’s the canonical source for x86 memory ordering rules.
Write-Back (WB) Memory (System RAM Default)
First, for standard WB memory (the default for most system RAM), lock-prefixed instructions like lock cmpxchg do enforce strong barrier semantics for normal memory accesses: per Volume 3, Section 8.2.2 of the SDM, regular reads and writes cannot be reordered across a lock-prefixed instruction. The manual states directly:
Reads or writes cannot be reordered with I/O instructions, lock-prefixed instructions, or serializing instructions.
But here’s the nuance with weakly-ordered (non-temporal) accesses: the same section includes explicit exceptions for weak-ordered stores. Specifically, the SDM outlines these core ordering rules (with key exceptions):
- Reads are never reordered with other reads
- Writes are never reordered with earlier reads
- Memory writes are never reordered with other writes except:
Stream stores (writes) executed using non-temporal move instructions (MOVNTI, MOVNTQ, MOVNTDQ, MOVNTPS, MOVNTPD); and
String operations (see Section 8.2.4.1).
You might notice that the rule about lock-prefixed instructions doesn’t call out non-temporal accesses as an exception—so does that mean lock acts as a full barrier even for MOVNT* operations? Not exactly. Other sections of the SDM explicitly clarify that when working with weakly-ordered instructions, you need mfence (for full memory ordering) or sfence (for store-store ordering) to enforce proper sequencing. These sections never mention lock-prefixed instructions as a valid alternative for fencing non-temporal accesses.
The practical takeaway for WB memory: Lock-prefixed instructions do act as a full barrier for normal accesses, but they are not a reliable substitute for mfence/sfence when dealing with non-temporal (weakly-ordered) operations. Even though the initial ordering rule doesn’t exclude non-temporal accesses, the SDM’s targeted guidance elsewhere makes it clear that lock prefixes aren’t designed to handle these weak-ordered operations.
Write-Combining (WC) Memory
For WC memory (used for specialized hardware like GPU framebuffers or memory-mapped I/O), the rules are even more restrictive. WC memory has inherent weak ordering, and lock-prefixed instructions offer little to no reliable barrier or atomicity guarantees here. The SDM explicitly notes that lock operations on WC memory don’t behave the same as on WB memory. To enforce ordering for WC accesses, you must use explicit fencing instructions like sfence or mfence—lock prefixes are not a valid replacement.
内容的提问来源于stack exchange,提问作者BeeOnRope

