混合不同大小原子操作的原子对象内存序及代码有效性问询
问题
更新说明
我已了解这在ISO C中属于未定义行为(UB),为之前表述模糊致歉。
本问题源自我此前的Stack Overflow问题(原标题:Can atomic operations of different sizes be mixed?)
假设该问题的结论成立(x86架构下成立,其他架构大概率成立但目前无法保证),请问以下代码是否合法有效?
#include <stdint.h> #include <assert.h> #include <stdatomic.h> #include <pthread.h> struct vec { uint64_t low; uint64_t high; }; union test { struct { _Atomic(uint64_t) low; _Atomic(uint64_t) high; }; _Atomic struct vec atomic; }; static union test vec0; static union test vec1; static void *writter0(void *arg) { for (uint64_t i = 0; i < UINT64_MAX; ++i) { atomic_store_explicit(&vec0.low, i, memory_order_relaxed); atomic_store_explicit(&vec0.high, i, memory_order_release); } return NULL; } static void *reader0(void *arg) { while (1) { const struct vec raw = atomic_load_explicit(&vec0.atomic, memory_order_relaxed); assert(raw.high <= raw.low); } return NULL; } static void *writter1(void *arg) { for (uint64_t i = 0; i < UINT64_MAX; ++i) { const struct vec raw = { .low = i, .high = i }; atomic_store_explicit(&vec1.atomic, raw, memory_order_relaxed); } return NULL; } static void *reader1(void *arg) { while (1) { const uint64_t low = atomic_load_explicit(&vec1.low, memory_order_acquire); const uint64_t high = atomic_load_explicit(&vec1.high, memory_order_relaxed); assert(high >= low); } return NULL; } int main(void) { pthread_t th[4]; assert(pthread_create(&th[0], NULL, writter0, NULL) == 0); assert(pthread_create(&th[1], NULL, reader0, NULL) == 0); assert(pthread_create(&th[2], NULL, writter1, NULL) == 0); assert(pthread_create(&th[3], NULL, reader1, NULL) == 0); assert(pthread_join(th[0], NULL) == 0); assert(pthread_join(th[1], NULL) == 0); assert(pthread_join(th[2], NULL) == 0); assert(pthread_join(th[3], NULL) == 0); return 0; }
上述代码使用了未配对的release/acquire操作,这点较为特殊。
疑问
- 基于x86架构下混合不同大小原子操作的结论成立的前提,这段代码是否合法有效?
- 若代码中的断言失败,会出现在哪些硬件架构及对应的指令下?
回答
代码合法性分析
首先明确:即使在x86架构下,这段代码依然不符合ISO C标准定义——标准明确禁止通过union的不同原子成员(_Atomic(uint64_t)和_Atomic struct vec)同时访问同一内存区域,这属于未定义行为。
但从x86硬件的实际运行逻辑来看:
- x86的内存模型是TSO(总顺序一致),自然对齐的单条加载/存储操作都是原子的,16字节结构体的原子加载用普通
mov指令,原子存储也用mov指令。 writter0与reader0组合:写入时按程序顺序先relaxed写low、再release写high;x86的TSO模型保证写操作对其他核可见的顺序与程序顺序一致,因此reader0读到的high值要么等于low(同一轮i的写入),要么小于low(未读到最新的high写入),断言raw.high <= raw.low不会触发。writter1与reader1组合:写入时原子存储整个16字节结构体;x86的acquire加载会阻止后续加载重排到它之前,因此acquire读low后,relaxed读high一定能看到low对应写入之后的high值,断言high >= low也不会触发。
不过这只是x86硬件的实际表现,不代表代码合法——标准未定义跨union原子成员的访问行为,编译器可能进行重排、合并等优化,导致代码行为超出预期。
断言可能失败的架构场景
断言失败的核心原因是硬件内存模型允许写操作重排,或不同大小原子操作的可见性不遵循预期顺序,典型场景如下:
ARM(v7及更早,v8部分弱内存模型场景)
ARM是弱内存模型,默认允许写操作重排,原子操作的内存序依赖显式屏障。
- 对于
writter0与reader0:若硬件允许high的release写重排到low的relaxed写之前,reader0可能读到high为i、low为i-1,触发断言失败。对应指令为普通存储str和16字节原子存储strd,无足够屏障时写操作可能乱序。 - 对于
writter1与reader1:若ARM将16字节原子存储拆分为两个独立的64位原子存储且无同步屏障,reader1的acquire读low后,relaxed读high可能读到旧值,导致断言失败。
PowerPC
PowerPC同样是弱内存模型,支持写-写、读-读重排(无屏障时)。
- 对于
writter0:PowerPC的普通存储std和带release语义的存储,若无sync或lwsync屏障,high的写可能先于low的写被其他核看到,导致reader0读到的high值大于low值。 - 对于
writter1:若16字节原子存储拆分为两个独立64位原子存储,reader1的acquire读low后,relaxed读high可能读到旧值,触发断言失败。对应指令包括std、lwarx/stwcx.等原子操作指令。
RISC-V(RVWMO弱有序内存模型配置)
RISC-V默认RVWMO模型允许写-写、读-写重排。
- 对于
writter0:low的relaxed写和high的release写可能被重排,reader0读到的high值大于low值,触发断言失败。对应指令为64位存储sd、原子存储amoswap.d,需fence rw,w类屏障保证release语义顺序。 - 对于
writter1:若16字节原子存储拆分为两个独立64位原子存储,reader1的acquire读low后,relaxed读high可能读到旧值,导致断言失败。
内容的提问来源于stack exchange,提问作者untitled
相关产品推荐
相关产品推荐

