C++多线程生成随机数为何与单线程效率相近?
多线程随机数生成性能问题排查与修复
我尝试编写一个多线程程序,生成包含N*NumPerThread个均匀随机整数的vector,其中N是std::thread::hardware_concurrency()的返回值,NumPerThread为每个线程要生成的随机数数量。但多线程版本和单线程版本执行时间大致相同,怀疑多线程版本存在问题,特此咨询。
多线程版本代码
#include <iostream> #include <thread> #include <vector> #include <random> #include <chrono> using Clock = std::chrono::high_resolution_clock; namespace Vars { const unsigned int N = std::thread::hardware_concurrency(); //number of threads on device const unsigned int NumPerThread = 5e5; //number of random numbers to generate per thread std::vector<int> RandNums(NumPerThread*N); std::random_device rd; std::mt19937 gen(rd()); std::uniform_int_distribution<> dis(1, 1000); int sz = 0; } using namespace Vars; void AddN(int start) { static std::mutex mtx; std::lock_guard<std::mutex> lock(mtx); for (unsigned int i=start; i<start+NumPerThread; i++) { RandNums[i] = dis(gen); ++sz; } } int main() { auto start_time = Clock::now(); std::vector<std::thread> threads; threads.reserve(N); for (unsigned int i=0; i<N; i++) { threads.emplace_back(std::move(std::thread(AddN, i*NumPerThread))); } for (auto &i: threads) { i.join(); } auto end_time = Clock::now(); std::cout << "\nTime difference = " << std::chrono::duration<double, std::nano>(end_time - start_time).count() << " nanoseconds\n"; std::cout << "size = " << sz << '\n'; }
单线程版本代码
#include <iostream> #include <thread> #include <vector> #include <random> #include <chrono> using Clock = std::chrono::high_resolution_clock; namespace Vars { const unsigned int N = std::thread::hardware_concurrency(); //number of threads on device const unsigned int NumPerThread = 5e5; //number of random numbers to generate per thread std::vector<int> RandNums(NumPerThread*N); std::random_device rd; std::mt19937 gen(rd()); std::uniform_int_distribution<> dis(1, 1000); int sz = 0; } using namespace Vars; void AddN() { for (unsigned int i=0; i<NumPerThread*N; i++) { RandNums[i] = dis(gen); ++sz; } } int main() { auto start_time = Clock::now(); AddN(); auto end_time = Clock::now(); std::cout << "\nTime difference = " << std::chrono::duration<double, std::nano>(end_time - start_time).count() << " nanoseconds\n"; std::cout << "size = " << sz << '\n'; }
问题分析
你的多线程版本完全没发挥并行优势,核心问题有两点:
- 全局互斥锁导致完全串行:
AddN函数里的std::lock_guard把整个for循环锁住,所有线程必须排队执行,和单线程逻辑完全一致,甚至因为线程切换开销可能更慢。 - 全局随机数生成器的竞争:
std::mt19937不是线程安全的,多线程同时调用会触发未定义行为;加锁后又彻底串行化,完全失去多线程意义。
修复方案
要实现真正的并行,需要做到:每个线程独立处理专属区间,使用独立的随机数生成器,避免不必要的同步。
修复后的多线程代码:
#include <iostream> #include <thread> #include <vector> #include <random> #include <chrono> using Clock = std::chrono::high_resolution_clock; namespace Vars { const unsigned int N = std::thread::hardware_concurrency(); const unsigned int NumPerThread = 500000; std::vector<int> RandNums(NumPerThread * N); } using namespace Vars; void AddN(int start) { // 每个线程拥有独立的随机数生成器,避免竞争 std::random_device rd; std::mt19937 gen(rd()); std::uniform_int_distribution<> dis(1, 1000); // 操作预先划分的专属索引区间,无资源竞争无需锁 for (unsigned int i = start; i < start + NumPerThread; ++i) { RandNums[i] = dis(gen); } } int main() { auto start_time = Clock::now(); std::vector<std::thread> threads; threads.reserve(N); for (unsigned int i = 0; i < N; ++i) { threads.emplace_back(AddN, i * NumPerThread); } for (auto& t : threads) { t.join(); } auto end_time = Clock::now(); std::cout << "\nTime difference = " << std::chrono::duration<double, std::nano>(end_time - start_time).count() << " nanoseconds\n"; std::cout << "size = " << RandNums.size() << '\n'; }
优化说明
- 每个线程创建独立的
std::mt19937和分布器,彻底避免线程间竞争,实现真正并行计算 - 移除全局互斥锁,因为线程操作的vector区间完全不重叠,无同步必要
- 用
RandNums.size()替代sz变量,省去线程累加的同步开销 - 将
5e5改为500000,避免浮点数转整数的潜在问题
内容的提问来源于stack exchange,提问作者Edward Finkelstein
相关产品推荐
相关产品推荐

