为何Glibc 2.32中x86读写内存屏障未使用__volatile asm?
__asm __volatile is used for full barrier but not read/write barriers in Glibc 2.32 x86? Great question! Let's unpack this by looking at x86's memory model, what each barrier does, and how GCC's __asm and __asm __volatile behave.
First, let's recall the x86 memory model quirk: it's a strongly ordered architecture by default. Here's what that means for memory operations:
- Loads (reads) won't be reordered with other loads (
LoadLoadordering is guaranteed) - Stores (writes) won't be reordered with other stores (
StoreStoreordering is guaranteed) - Loads won't be reordered with earlier stores (
LoadStoreordering is guaranteed) - The only reordering allowed is stores being reordered after later loads (
StoreLoadreordering)
Breaking down each barrier definition:
atomic_read_barrier()andatomic_write_barrier()#define atomic_read_barrier() __asm ("" ::: "memory") #define atomic_write_barrier() __asm ("" ::: "memory")These are compiler-only memory barriers. The empty
__asmblock with the"memory"constraint tells GCC:- Don't reorder memory accesses across this barrier
- Flush any cached memory values in registers to main memory, or reload values from memory as needed
Since x86 hardware already enforces the ordering required for read/write barriers (no LoadLoad/StoreStore/LoadStore reordering), we don't need any actual machine instructions here. The
__volatilemodifier isn't necessary because:- The
"memory"constraint already forces the compiler to retain this barrier (it can't optimize it away, as it affects memory operation ordering) - There's no actual hardware instruction to optimize out—this is purely a compiler hint.
atomic_full_barrier()#define atomic_full_barrier() \ __asm __volatile (LOCK_PREFIX "orl $0, (%%" SP_REG ")" ::: "memory")A full memory barrier needs to prevent all memory reordering, including the only allowed one on x86:
StoreLoadreordering. To do this, we need a hardware-level barrier.The
LOCK_PREFIX "orl $0, (%%rsp)"is a trick to trigger a full memory barrier:orl $0, (%rsp)is a no-op (it doesn't change the value at the stack pointer)- But the
LOCKprefix forces the CPU to issue a memory fence, which enforces all memory ordering rules.
Here's why
__volatileis required:- Without it, the compiler might look at this
orl $0instruction and think it's a useless no-op with no side effects. It could optimize the entire__asmblock away, which would remove the hardware barrier entirely. __volatiletells the compiler: "Don't touch this assembly block—even if it looks like it does nothing, it has important hardware-level side effects."
To sum it up:
- Read/write barriers only need to constrain the compiler's optimizations, which the
"memory"constraint in a plain__asmblock handles perfectly. - The full barrier relies on a hardware instruction that the compiler might otherwise optimize out, so
__volatileis needed to force the compiler to emit it.
内容的提问来源于stack exchange,提问作者Zihe Liu

