跨循环迭代时,relaxed存储能否重排至compare_exchange之前?
代码示例
#include <atomic> #include <cassert> #include <thread> struct foo { std::atomic<foo*> _next = nullptr; }; foo dummy = {}; foo a = {}; std::atomic<foo*> b = nullptr; void thread_0() { auto* old = b.load(std::memory_order_acquire); // (1) do { a._next.store(old, std::memory_order_relaxed); // First iteration: (2), Second iteration: (4) } while (!b.compare_exchange_weak(old, &a, std::memory_order_release, std::memory_order_relaxed)); // First iteration: (3), Second iteration: (5) } void thread_1() { // plain store is enough for the demonstration here, // but then the code may hang indefinitely when it is actually executed. // Just think this is a plain store in this example. const auto* avoid_hang = b.exchange(&dummy, std::memory_order_relaxed); // (6) if (avoid_hang) { return; } while (b.load(std::memory_order_acquire) != &a) {}; // (7) assert(a._next.load(std::memory_order_relaxed)); // Can this assert fire? } /* void thread_1_ignore_hang() { b.store(&dummy, std::memory_order_relaxed); // (6) while (b.load(std::memory_order_acquire) != &a) {}; // (7) assert(a._next.load(std::memory_order_relaxed)); // Can this assert fire? } */ int main() { std::jthread t0(thread_0); std::jthread t1(thread_1); return 0; }
指定执行顺序
Thread 0 Thread 1 (1) old = nullptr(load b, mo_acquire) (2) a->_next = nullptr (6) b = &dummy (3) cmpxchg fails, old = &dummy (load b, mo_relaxed) (4) a->_next = &dummy (5) cmpxchg succeeds, b = &a (load b, mo_relaxed, store to b, mo_release) (7) (load b, mo_acquire)
问题
已知std::memory_order_release存储(由compare_exchange_weak引入)仅会阻止其之前的读写操作被重排至其后。现咨询:在thread_1视角下,操作(4)是否可能被重排至操作(2)之前,导致读取a._next为nullptr并触发断言?因x86硬件原子操作默认带acquire/release语义,无法测试验证。
回答
这个断言不会触发,原因如下:
线程内部的执行顺序保障:
操作(2)和(4)属于同一个do-while循环的两次不同迭代,线程内部的指令重排必须遵循「as-if规则」——即重排后的执行效果必须和程序串行执行的效果一致。在thread_0的串行执行流中,必然是先完成第一次迭代的(2)、(3),才会进入第二次迭代执行(4)、(5)。因此即使是memory_order_relaxed的存储操作,编译器和CPU也不会把第二次迭代的(4)重排到第一次迭代的(2)之前,破坏循环的串行执行逻辑。Release-Acquire同步关系的保障:
thread_0中(5)的compare_exchange_weak成功时,使用了memory_order_release语义存储&a到b;thread_1中(7)使用memory_order_acquire语义加载b得到&a,这两者形成了release-acquire同步关系。根据C++内存模型,这种同步关系确保thread_0中(5)之前的所有内存操作(包括(4)对a._next的存储),对执行(7)后的thread_1完全可见。
因此,当thread_1执行到断言时,必然能看到(4)写入的&dummy,断言不会触发。
内容的提问来源于stack exchange,提问作者whatishappened

