You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

C++多线程生成随机数为何与单线程效率相近?

多线程随机数生成性能问题排查与修复

我尝试编写一个多线程程序,生成包含N*NumPerThread个均匀随机整数的vector,其中N是std::thread::hardware_concurrency()的返回值,NumPerThread为每个线程要生成的随机数数量。但多线程版本和单线程版本执行时间大致相同,怀疑多线程版本存在问题,特此咨询。

多线程版本代码

#include <iostream>
#include <thread>
#include <vector>
#include <random>
#include <chrono>

using Clock = std::chrono::high_resolution_clock;

namespace Vars
{
    const unsigned int N = std::thread::hardware_concurrency(); //number of threads on device
    const unsigned int NumPerThread = 5e5; //number of random numbers to generate per thread
    std::vector<int> RandNums(NumPerThread*N);
    std::random_device rd;
    std::mt19937 gen(rd());
    std::uniform_int_distribution<> dis(1, 1000);
    int sz = 0;
}

using namespace Vars;

void AddN(int start)
{
    static std::mutex mtx;
    std::lock_guard<std::mutex> lock(mtx);
    for (unsigned int i=start; i<start+NumPerThread; i++)
    {
        RandNums[i] = dis(gen);
        ++sz;
    }
}

int main()
{
    auto start_time = Clock::now();
    std::vector<std::thread> threads;
    threads.reserve(N);
    
    for (unsigned int i=0; i<N; i++)
    {
        threads.emplace_back(std::move(std::thread(AddN, i*NumPerThread)));
    }

    for (auto &i: threads)
    {
        i.join();
    }
        
    auto end_time = Clock::now();
    std::cout << "\nTime difference = "
    << std::chrono::duration<double, std::nano>(end_time - start_time).count() << " nanoseconds\n";
    std::cout << "size = " << sz << '\n';
}

单线程版本代码

#include <iostream>
#include <thread>
#include <vector>
#include <random>
#include <chrono>


using Clock = std::chrono::high_resolution_clock;



namespace Vars
{
    const unsigned int N = std::thread::hardware_concurrency(); //number of threads on device
    const unsigned int NumPerThread = 5e5; //number of random numbers to generate per thread
    std::vector<int> RandNums(NumPerThread*N);
    std::random_device rd;
    std::mt19937 gen(rd());
    std::uniform_int_distribution<> dis(1, 1000);
    int sz = 0;
}
    


using namespace Vars;


void AddN()
{
    for (unsigned int i=0; i<NumPerThread*N; i++)
    {
        RandNums[i] = dis(gen);
        ++sz;
    }
}

int main()
{
    auto start_time = Clock::now();

    AddN();
    
    auto end_time = Clock::now();
    std::cout << "\nTime difference = "
    << std::chrono::duration<double, std::nano>(end_time - start_time).count() << " nanoseconds\n";
    std::cout << "size = " << sz << '\n';
}

问题分析

你的多线程版本完全没发挥并行优势,核心问题有两点:

  1. 全局互斥锁导致完全串行:AddN函数里的std::lock_guard把整个for循环锁住,所有线程必须排队执行,和单线程逻辑完全一致,甚至因为线程切换开销可能更慢。
  2. 全局随机数生成器的竞争:std::mt19937不是线程安全的,多线程同时调用会触发未定义行为;加锁后又彻底串行化,完全失去多线程意义。

修复方案

要实现真正的并行,需要做到:每个线程独立处理专属区间,使用独立的随机数生成器,避免不必要的同步。

修复后的多线程代码:

#include <iostream>
#include <thread>
#include <vector>
#include <random>
#include <chrono>

using Clock = std::chrono::high_resolution_clock;

namespace Vars
{
    const unsigned int N = std::thread::hardware_concurrency();
    const unsigned int NumPerThread = 500000;
    std::vector<int> RandNums(NumPerThread * N);
}

using namespace Vars;

void AddN(int start)
{
    // 每个线程拥有独立的随机数生成器,避免竞争
    std::random_device rd;
    std::mt19937 gen(rd());
    std::uniform_int_distribution<> dis(1, 1000);

    // 操作预先划分的专属索引区间,无资源竞争无需锁
    for (unsigned int i = start; i < start + NumPerThread; ++i)
    {
        RandNums[i] = dis(gen);
    }
}

int main()
{
    auto start_time = Clock::now();
    std::vector<std::thread> threads;
    threads.reserve(N);
    
    for (unsigned int i = 0; i < N; ++i)
    {
        threads.emplace_back(AddN, i * NumPerThread);
    }

    for (auto& t : threads)
    {
        t.join();
    }
        
    auto end_time = Clock::now();
    std::cout << "\nTime difference = "
              << std::chrono::duration<double, std::nano>(end_time - start_time).count() << " nanoseconds\n";
    std::cout << "size = " << RandNums.size() << '\n';
}

优化说明

  • 每个线程创建独立的std::mt19937和分布器,彻底避免线程间竞争,实现真正并行计算
  • 移除全局互斥锁,因为线程操作的vector区间完全不重叠,无同步必要
  • 用RandNums.size()替代sz变量,省去线程累加的同步开销
  • 将5e5改为500000,避免浮点数转整数的潜在问题

内容的提问来源于stack exchange,提问作者Edward Finkelstein

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.09 17:15:33