C++标准中RMW操作的Load部分是否比纯Load要求更高?附实例分析
首先看以下代码实例:
#include <iostream> #include <atomic> #include <thread> #include <chrono> #include <cassert> int main(){ std::atomic<int> v = 0; std::atomic<bool> flag = false; std::thread t1([&](){ while(!flag.load(std::memory_order::relaxed)){} // #1 assert(v.exchange(2,std::memory_order::relaxed) == 1); // #2 }); std::thread t2([&](){ if(v.exchange(1,std::memory_order::relaxed) == 0){ // #3 flag.store(true, std::memory_order::relaxed); // #4 } }); t1.join(); t2.join(); }
在该实例中,#1处的循环仅在#4将flag设为true时退出,而flag被设为true的前提是#3处RMW操作的Load部分读取到值0。根据C++标准[atomics.order] p10的规则:
Atomic read-modify-write operations shall always read the last value (in the modification order) written before the write associated with the read-modify-write operation.
这意味着如果#3处的RMW操作读取到0,那么其他RMW操作就无法再读取到0。换句话说,如果#2能读取到0并写入2,那#3就不会读取到0,#4也不会执行——这正是自旋锁的核心工作原理:当所有操作都是RMW操作时,每个RMW操作读取的值是独有的。
基于此提出三个问题:
Q1:#2处的断言是否永远不会触发失败?
如果把#2改为纯Load操作,代码如下:
assert(v.load(std::memory_order::relaxed) == 1); // #2'
根据C++标准[intro.races] p18的规则:
If a side effect X on an atomic object M happens before a value computation B of M, then the evaluation B takes its value from X or from a side effect Y that follows X in the modification order of M.
在#2'之前发生的副作用只有初始值0,尽管#3处写入的1在修改顺序中位于0之后,但纯Load操作仍有可能读取到0。这一点也能在[atomics.order] p11的建议实践中得到印证:
Recommended practice: The implementation should make atomic stores visible to atomic loads, and atomic loads should observe atomic stores, within a reasonable amount of time.
从实现角度看,存在时间延迟可能导致#3的存储在合理时间内对#2'不可见;从C++标准的角度看,这也由[intro.races] p18中的“or”所隐含。
Q2:若将#2处的RMW操作改为#2'这样的纯Load操作,断言是否可能触发失败?
Q3:若#2'可能失败而#2永远不会失败,是否意味着至少在该实例中,RMW操作比纯Load操作更倾向于读取修改顺序中较晚的修改值?
补充说明:
我认为编译器不会对#2和#1进行重排序,因为#2类似自旋锁中的失败CAS(即采用std::memory_order::relaxed内存序的纯Load操作),如果存在重排序,自旋锁将无法正常工作。此外,本实例中编译器的任何重排序都会被断言观测到。不过从内存序角度看,#3与#2之间不存在先行关系,理论上#3可能失败,对此我不确定。但本实例依赖执行的逻辑顺序,任何对该顺序的破坏都可被观测到。
内容的提问来源于stack exchange,提问作者xmh0511

