原子变量更新跨线程反射延迟的测量与疑问
跨线程原子变量写入的感知延迟探究
实验设计
为探究变量写入操作在跨线程间被感知的最小延迟,采用全局atomic<int64_t>变量,由一个线程周期性更新,另一个线程自旋检测更新值,两个线程绑定到Ubuntu系统的独立核心。
全局变量定义
// global constexpr int total = 100; atomic<int64_t> var;
读取线程代码
void reader() { int count = 0; int64_t tps[total]; int64_t last = 0; while(count < total) { int64_t send_tp = var.load(std::memory_order_seq_cst); auto tp = high_resolution_clock::now(); int64_t curr = duration_cast<nanoseconds>(tp.time_since_epoch()).count(); if (send_tp != last) { last = send_tp; tps[count] = curr - send_tp; count++; } } for(auto i = 0; i<total; i++) cout << tps[i] << endl; }
写入线程代码
void writer() { for (int i=0; i<total; i++) { auto tp = high_resolution_clock::now(); int64_t curr = duration_cast<nanoseconds>(tp.time_since_epoch()).count(); var.store(curr, std::memory_order_seq_cst); // 添加写入间隔,避免读取线程遗漏更新 while(duration_cast<nanoseconds>(high_resolution_clock::now() - tp).count() < 100000000); } }
单线程原子操作开销测量代码
为排除原子操作本身的开销影响,单独测量单线程下的原子操作开销:
void overhead() { int count = 0; int64_t tps[total]; int64_t last = 0; while(count < total) { auto tp1 = high_resolution_clock::now(); int64_t to_send = duration_cast<nanoseconds>(tp1.time_since_epoch()).count(); var.store(to_send, std::memory_order_seq_cst); int64_t send_tp = var.load(std::memory_order_seq_cst); auto tp = high_resolution_clock::now(); int64_t curr = duration_cast<nanoseconds>(tp.time_since_epoch()).count(); if (send_tp != last) { last = send_tp; tps[count] = curr - send_tp; count++; } } for(auto i = 0; i<total; i++) cout << tps[i] << endl; }
实验结果
- 跨线程场景下,测得中位数延迟约70纳秒
- 单线程原子操作的中位数开销约30纳秒(推测主要来自
chrono::high_resolution_clock()的调用开销) - 由此推算跨线程的实际延迟约40纳秒
- 尝试
memory_order_relaxed、release-acquire等不同内存序,结果差异不大
疑问与优化方向
按理论理解,跨线程同步仅需从相邻核心获取L1缓存行,为何测得延迟约40纳秒?实验中是否遗漏了关键因素?有哪些优化方向可以更精准测量跨线程感知延迟?
硬件与编译信息
- 硬件:Intel(R) Core(TM) i9-9900K CPU(超线程已禁用)
- 编译命令:
g++ file.cpp -lpthread -O3
内容的提问来源于stack exchange,提问作者W1nTer003
相关产品推荐
相关产品推荐

