Windows平台std::shared_mutex无排他锁时unlock_shared()仍阻塞的死锁问题
Windows平台SRW锁疑似死锁Bug分析与复现
问题背景
我们团队遇到一个死锁问题,怀疑是Windows平台SRW锁实现存在Bug。
复现逻辑
- 主线程获取排他锁
- 主线程创建N个子线程
- 每个子线程执行:
- 获取共享锁
- 自旋等待所有子线程都获取到共享锁
- 释放共享锁
- 主线程释放排他锁
注:此逻辑可通过C++20的
std::latch实现,但这并非重点。
死锁现象
该代码多数情况下正常运行,但每约5000次循环会出现一次死锁。死锁时,恰好1个子线程成功获取共享锁,其余N-1个子线程卡在lock_shared()中。在Windows系统中,该函数会调用RtlAcquireSRWLockShared,并在NtWaitForAlertByThreadId中阻塞。
无论直接使用std::shared_mutex、std::shared_lock/std::unique_lock,还是直接调用SRW相关函数,均会出现此现象。
问题分析
2017年Raymond Chen的一篇博客曾提及该现象,但当时被归咎于用户错误。我们认为这是SRW锁的Bug:若子线程不等待直接调用unlock_shared(),则会唤醒阻塞的兄弟线程。而std::shared_lock或SRW相关文档中均未说明无活跃排他锁时允许阻塞。
该问题未在非Windows平台出现。
复现代码
#include <atomic> #include <cstdint> #include <iostream> #include <memory> #include <shared_mutex> #include <thread> #include <vector> struct ThreadTestData { int32_t numThreads = 0; std::shared_mutex sharedMutex = {}; std::atomic<int32_t> readCounter; }; int DoStuff(ThreadTestData* data) { // Acquire reader lock data->sharedMutex.lock_shared(); // wait until all read threads have acquired their shared lock data->readCounter.fetch_add(1); while (data->readCounter.load() != data->numThreads) { std::this_thread::yield(); } // Release reader lock data->sharedMutex.unlock_shared(); return 0; } int main() { int count = 0; while (true) { ThreadTestData data = {}; data.numThreads = 5; // Acquire write lock data.sharedMutex.lock(); // Create N threads std::vector<std::unique_ptr<std::thread>> readerThreads; readerThreads.reserve(data.numThreads); for (int i = 0; i < data.numThreads; ++i) { readerThreads.emplace_back(std::make_unique<std::thread>(DoStuff, &data)); } // Release write lock data.sharedMutex.unlock(); // Wait for all readers to succeed for (auto& thread : readerThreads) { thread->join(); } // Cleanup readerThreads.clear(); // Spew so we can tell when it's deadlocked count += 1; std::cout << count << std::endl; } return 0; }
并行栈情况
并行栈截图显示:主线程正常阻塞在thread::join()上,一个子线程已获取锁并处于yield循环,四个子线程阻塞在lock_shared()中。
内容的提问来源于stack exchange,提问作者LordCecil
相关产品推荐
相关产品推荐

