You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

std::shared_timed_mutex::try_lock_until触发ThreadSanitizer数据竞争问题排查

问题背景

我正在编写测试用例验证std::shared_timed_mutex::try_lock_until的行为。

初始代码

#include <thread>
#include <iostream>
#include <chrono>
#include <shared_mutex>
#include <cassert>
 
std::shared_timed_mutex test_mutex;
int global;
 
void f()
{
    auto now=std::chrono::steady_clock::now();
    test_mutex.try_lock_until(now + std::chrono::seconds(100));
    //test_mutex.lock();
    --global;
    std::cout << "In lock, global=" << global << '\n';
    test_mutex.unlock();
}

void g()
{
    auto now=std::chrono::steady_clock::now();
    test_mutex.try_lock_shared_until(now + std::chrono::seconds(10));
    //test_mutex.lock_shared();
    std::cout << "In shared lock, global=" << global << '\n';
    test_mutex.unlock_shared();
}
 
int main()
{
    global = 1;
    test_mutex.lock_shared();
    std::thread t1(f);
    std::thread t2(g);
    test_mutex.unlock_shared();
    t1.join();
    t2.join();
    assert(global == 0);
}

预期执行流程

  1. main函数先获取读锁,再启动线程f和g
  2. f尝试获取独占锁进入阻塞状态
  3. g获取读锁,读取global变量后释放读锁
  4. main函数释放持有的读锁
  5. f解除阻塞,修改global变量后释放锁结束运行
  6. 主线程等待f和g执行完成,断言验证global等于0后退出
  • 第2、3步执行顺序可互换

问题现象

该代码单独运行表现正常,用gdb在g的global读取操作和f的global写入操作处打断点,运行后会先停在读取断点处,符合预期。但使用-fsanitize=thread参数编译后,ThreadSanitizer报出数据竞争告警,告警日志如下:

WARNING: ThreadSanitizer: data race (pid=6780)
  Read of size 4 at 0x000000407298 by thread T2:
    #0 g() /home/paulf/scratch/valgrind/drd/tests/try_lock_shared_until14.cpp:25 (try_lock_shared_until14+0x402484)
[trimmed]
    #6 execute_native_thread_routine ../../../../../libstdc++-v3/src/c++11/thread.cc:82 (libstdc++.so.6+0xd9c83)

  Previous write of size 4 at 0x000000407298 by thread T1:
    #0 f() /home/paulf/scratch/valgrind/drd/tests/try_lock_shared_until14.cpp:15 [triimed]
    #6 execute_native_thread_routine ../../../../../libstdc++-v3/src/c++11/thread.cc:82 (libstdc++.so.6+0xd9c83)

  Location is global 'global' of size 4 at 0x000000407298 (try_lock_shared_until14+0x000000407298)

用gdb调试tsan编译后的版本时,发现f不会在独占锁处阻塞,会先执行写入操作。我清楚当前示例写法不严谨,应该检查锁接口的返回值且不能依赖超时逻辑。我想知道ThreadSanitizer到底改变了什么行为?如果我换成普通的lock/lock_shared/unlock/unlock_shared接口,ThreadSanitizer就不会报数据竞争。

  • 注:我无法使用DRD或Helgrind复现该问题,我正在为这两个工具编写测试用例,我使用的Fedora 34 + GCC 11.2.1 amd64平台目前暂不支持相关特性。

修复后的代码

我已更新到第三版可正常运行的代码,主线程通过条件变量等待g执行完成后再释放持有的共享锁,之后f才能获取到独占锁,修复后的代码如下:

#include <thread>
#include <iostream>
#include <chrono>
#include <shared_mutex>
#include <mutex>
#include <cassert>
#include <condition_variable>

std::shared_timed_mutex test_mutex;
std::mutex cv_mutex;
std::condition_variable cv;
int global;
bool reads_done = false;
 
void f()
{
    auto now=std::chrono::steady_clock::now();
    std::cout << "In lock, trying to get mutex\n";
    if (test_mutex.try_lock_until(now + std::chrono::seconds(3)))
    {
       --global;
       std::cout << "In lock, global=" << global << '\n';
       test_mutex.unlock();
    }
    else
    {
        std::cerr << "Lock failed\n";
    }
}

void g()
{
    auto now=std::chrono::steady_clock::now();
    std::cout << "In shared lock, trying to get mutex\n";
    if (test_mutex.try_lock_shared_until(now + std::chrono::seconds(2)))
    {
       std::cout << "In shared lock, global=" << global << '\n';
       test_mutex.unlock_shared();
    }
    else
    {
        std::cerr << "Lock shared failed\n";
    }
    std::unique_lock<std::mutex> lock(cv_mutex);
    reads_done = true;
    cv.notify_all();
}
 
int main()
{
    global = 1;
    test_mutex.lock_shared();
    std::thread t1(f);
    std::thread t2(g);
    {
       std::unique_lock<std::mutex> lock(cv_mutex);
       while (!reads_done)
       {
          cv.wait(lock);
       }
    }
    std::cout << "Main, reader thread done\n";
    test_mutex.unlock_shared();
    std::cout << "Main, no more shared locks\n";
    t1.join();
    t2.join();
    assert(global == 0);
}

原因说明
  1. ThreadSanitizer为了检测数据竞争,会故意调整线程调度时序、插入随机延迟,放大低概率的并发问题。原代码的正常执行依赖「g比main释放锁更早抢到共享锁」的时序,这个时序在普通运行下大概率发生,但ThreadSanitizer的调度打乱后,可能出现main先释放了自己持有的共享锁,此时f的独占锁请求已经满足,直接抢到独占锁先完成写入,之后g才尝试抢共享锁的情况,自然就出现了读写数据竞争。
  2. GCC 11版本的ThreadSanitizer对try_lock_until/try_lock_shared_until这类超时锁的同步关系建模存在局限性,无法准确识别超时等待的happens-before关系,也会提升问题暴露的概率。而普通lock接口的同步关系建模非常成熟,能正确识别独占锁和共享锁的互斥规则,因此不会报数据竞争。
  3. 原代码没有检查try_lock_*的返回值,哪怕锁超时失败,也会直接操作全局变量并调用unlock,这本身就属于未定义行为,ThreadSanitizer的调度只是把这个问题暴露了出来。

你更新的第三版代码通过条件变量保证了g执行完成后main才释放自己持有的共享锁,从根源上避免了f先抢到独占锁的可能,同时新增了锁获取成功的判断,消除了未定义行为,因此不会再触发数据竞争告警。


内容的提问来源于stack exchange,提问作者Paul Floyd

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.24 15:15:01