C++11顺序一致性与GCC__sync_synchronize的映射及等价性问询
一、旧GCC内置函数到C++11内存模型的映射
__sync_synchronize()与std::atomic_thread_fence(memory_order_seq_cst)的关系
- GCC的
__sync_synchronize()是全内存屏障,强制所有前置的加载/存储操作完成后,才执行后置的加载/存储操作,覆盖LoadLoad、LoadStore、StoreLoad、StoreStore四类屏障的语义。 - 从形式语义上看,
std::atomic_thread_fence(memory_order_seq_cst)不仅要求全屏障的内存约束,还会参与全局的SeqCst操作总序——所有SeqCst栅栏和SeqCst原子操作在全局范围内有统一的执行顺序。而__sync_synchronize()仅保证自身的屏障约束,没有明确绑定到这个全局总序。不过在实际商用CPU上,比如x86的mfence、PowerPC的hwsync,这类指令本身就同时满足全屏障和全局序的要求,所以两者在多数平台上表现等价。 - 结论:形式上
std::atomic_thread_fence(memory_order_seq_cst)的语义更强(绑定全局总序),但在绝大多数实际硬件上,两者生成的指令完全一致,行为无差异。
二、memory_order_relaxed的语义与三个实验的断言风险
memory_order_relaxed的原子操作不建立任何synchronize-with或happens-before关系,仅保证操作本身的原子性,不对周围的内存操作做任何约束。因此三个实验中的断言都存在失败的可能:
实验1:使用C11 atomic_thread_fence
// global static atomic_bool lock = false; static atomic_bool critical_section = false; // thread 1 atomic_store_explicit(&critical_section, true, memory_order_relaxed); atomic_thread_fence(memory_order_seq_cst); atomic_store_explicit(&lock, true, memory_order_relaxed); // thread 2 if (atomic_load_explicit(&lock, memory_order_relaxed)) { // We should really `memory_order_acquire` the `lock` // or `atomic_thread_fence(memory_order_acquire)` here, // or this assertion may fail, no? assert(atomic_load_explicit(&critical_section, memory_order_relaxed)); }
线程1的SeqCst栅栏只保证自身前后的内存操作顺序,但线程2用relaxed加载lock后,没有对应的acquire语义约束——线程2的critical_section加载可能被重排到lock加载之前,或者看不到线程1中critical_section的写入(因为没有synchronize-with关系),所以断言可能失败。
实验2:直接对原子存储使用SeqCst
// global static atomic_bool lock = false; static atomic_bool critical_section = false; // thread 1 atomic_store_explicit(&critical_section, true, memory_order_relaxed); atomic_store_explicit(&lock, true, memory_order_seq_cst); // thread 2 if (atomic_load_explicit(&lock, memory_order_relaxed)) { // Again we should really `memory_order_acquire` the `lock` // or `atomic_thread_fence(memory_order_acquire)` here, // or this assertion may fail, no? assert(atomic_load_explicit(&critical_section, memory_order_relaxed)); }
线程1的lock存储是SeqCst,但线程2用relaxed加载lock,无法触发synchronize-with关系——SeqCst的全局序只对SeqCst操作生效,线程2的relaxed加载不参与这个总序,所以依然无法保证看到critical_section的写入,断言可能失败。
实验3:使用GCC内置__sync_synchronize()
// global static atomic_bool lock = false; static atomic_bool critical_section = false; // thread 1 atomic_store_explicit(&critical_section, true, memory_order_relaxed); __sync_synchronize(); atomic_store_explicit(&lock, true, memory_order_relaxed); // thread 2 if (atomic_load_explicit(&lock, memory_order_relaxed)) { // we should somehow put a `LoadLoad` memory barrier here, // or the assert might fail, no? assert(atomic_load_explicit(&critical_section, memory_order_relaxed)); }
线程1的全屏障保证了critical_section写入在lock写入之前,但线程2的relaxed加载lock后,没有LoadLoad屏障或acquire语义约束——线程2的CPU可能提前加载critical_section(甚至缓存了旧值),导致看不到线程1的写入,断言可能失败。
为什么RPi 5上没出现断言失败?
RPi 5基于ARMv9架构,ARM的内存模型虽然是弱序,但实际硬件中某些操作的重排概率较低,且你的测试可能没有触发足够的并发压力或重排条件。弱序架构的重排是允许但不必然发生的,单次或少量测试很难复现,需要大量并发迭代或特定的测试用例才能触发。
内容的提问来源于stack exchange,提问作者usb_naming_is_confusing

