You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

原子变量更新跨线程反射延迟的测量与疑问

跨线程原子变量写入的感知延迟探究

实验设计

为探究变量写入操作在跨线程间被感知的最小延迟,采用全局atomic<int64_t>变量,由一个线程周期性更新,另一个线程自旋检测更新值,两个线程绑定到Ubuntu系统的独立核心。

全局变量定义

// global
constexpr int total = 100;
atomic<int64_t> var;

读取线程代码

void reader()
{
    int count = 0;
    int64_t tps[total];

    int64_t last = 0;
    while(count < total)
    {
        int64_t send_tp = var.load(std::memory_order_seq_cst);
        auto tp = high_resolution_clock::now();
        int64_t curr = duration_cast<nanoseconds>(tp.time_since_epoch()).count();    

        if (send_tp != last)
        {
            last = send_tp;
            tps[count] = curr - send_tp;
            count++;
        }
    }

    for(auto i = 0; i<total; i++)
        cout << tps[i] << endl;
}

写入线程代码

void writer()
{
    for (int i=0; i<total; i++)
    {
        auto tp = high_resolution_clock::now();
        int64_t curr = duration_cast<nanoseconds>(tp.time_since_epoch()).count();
        var.store(curr, std::memory_order_seq_cst);

        // 添加写入间隔,避免读取线程遗漏更新
        while(duration_cast<nanoseconds>(high_resolution_clock::now() - tp).count() < 100000000);
    }
}

单线程原子操作开销测量代码

为排除原子操作本身的开销影响,单独测量单线程下的原子操作开销:

void overhead() {
    int count = 0;
    int64_t tps[total];

    int64_t last = 0;
    while(count < total)
    {
        auto tp1 = high_resolution_clock::now();
        int64_t to_send = duration_cast<nanoseconds>(tp1.time_since_epoch()).count();
        var.store(to_send, std::memory_order_seq_cst);

        int64_t send_tp = var.load(std::memory_order_seq_cst);
        auto tp = high_resolution_clock::now();
        int64_t curr = duration_cast<nanoseconds>(tp.time_since_epoch()).count();    

        if (send_tp != last)
        {
            last = send_tp;
            tps[count] = curr - send_tp;
            count++;
        }
    }

    for(auto i = 0; i<total; i++)
        cout << tps[i] << endl;
}

实验结果

  • 跨线程场景下,测得中位数延迟约70纳秒
  • 单线程原子操作的中位数开销约30纳秒(推测主要来自chrono::high_resolution_clock()的调用开销)
  • 由此推算跨线程的实际延迟约40纳秒
  • 尝试memory_order_relaxed、release-acquire等不同内存序,结果差异不大

疑问与优化方向

按理论理解,跨线程同步仅需从相邻核心获取L1缓存行,为何测得延迟约40纳秒?实验中是否遗漏了关键因素?有哪些优化方向可以更精准测量跨线程感知延迟?

硬件与编译信息

  • 硬件:Intel(R) Core(TM) i9-9900K CPU(超线程已禁用)
  • 编译命令:g++ file.cpp -lpthread -O3

内容的提问来源于stack exchange,提问作者W1nTer003

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.21 07:00:08