如何禁用Clang对thread_local变量的表达式消除优化?
问题背景与现象
代码示例
thread_local int* tls = nullptr; // 使用libcontext切换栈 void jump_stack(); void* test() { // 调用jump_stack前,假设处于线程1 int *cur_tls = tls; jump_stack(); // 切换栈后,处于线程2 // 需要重新加载tls cur_tls = tls; }
运行环境
- 操作系统:Darwin Kernel Version 22.1.0(Apple M1芯片)
- Clang版本:Apple clang version 14.0.0 (clang-1400.0.29.202)
编译命令
clang++ -c test.cpp --std=c++11 -g -O0
生成的汇编代码
; void* test() { 0: ff c3 00 d1 sub sp, sp, #48 4: fd 7b 02 a9 stp x29, x30, [sp, #32] 8: fd 83 00 91 add x29, sp, #32 c: 00 00 00 90 adrp x0, 0x0 <ltmp0+0xc> 10: 00 00 40 f9 ldr x0, [x0] 14: 08 00 40 f9 ldr x8, [x0] 18: 00 01 3f d6 blr x8 1c: e0 07 00 f9 str x0, [sp, #8] ; int *cur_tls = tls; 20: 08 00 40 f9 ldr x8, [x0] 24: e8 0b 00 f9 str x8, [sp, #16] ; jump_stack(); 28: 00 00 00 94 bl 0x28 <ltmp0+0x28> 2c: e0 07 40 f9 ldr x0, [sp, #8] ; cur_tls = tls; 30: 08 00 40 f9 ldr x8, [x0] 34: e8 0b 00 f9 str x8, [sp, #16] ; } 38: a0 83 5f f8 ldur x0, [x29, #-8] 3c: fd 7b 42 a9 ldp x29, x30, [sp, #32] 40: ff c3 00 91 add sp, sp, #48 44: c0 03 5f d6 ret
编译后生成的汇编显示:调用jump_stack()前,tls的地址被缓存到栈的[sp, #16]位置;切换栈后重新赋值cur_tls = tls时,编译器直接复用了缓存的地址,导致获取的是线程1的tls实例,而非切换后的线程2实例。
解决方法
1. 用volatile修饰thread_local变量
修改变量声明,添加volatile关键字:
volatile thread_local int* tls = nullptr;
volatile会强制编译器每次访问该变量时都直接从内存读取,不会复用之前缓存的值,确保切换线程后能获取当前线程的tls实例。这是标准C++的解决方案,兼容性最好。
2. 标记jump_stack()存在未感知的副作用
给jump_stack()添加__attribute__((side_effect))属性,告诉编译器该函数会产生编译器无法检测的副作用(比如切换线程上下文):
__attribute__((side_effect)) void jump_stack();
这样编译器会在调用jump_stack()后,重新加载所有可能被副作用影响的变量,包括thread_local变量。
3. 使用编译选项禁用TLS缓存优化
在编译时添加-fno-tls-direct-seg-refs选项:
clang++ -c test.cpp --std=c++11 -g -O0 -fno-tls-direct-seg-refs
该选项会禁用ARM64平台上直接引用TLS段的优化,强制每次访问thread_local变量时都通过线程控制块重新加载,保证获取当前线程的实例。
内容的提问来源于stack exchange,提问作者ehds
相关产品推荐
相关产品推荐

