You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用std::async()多线程插入vector,性能远差于单线程求排查

问题分析与解决方案

你的异步版本性能大幅落后的核心原因集中在这几点:

1. 线程创建与调度开销远超任务本身

你给每个元素都单独创建一个线程,std::launch::async会直接启动新线程(主流实现逻辑)。线程的创建、调度、销毁都有明显的内核级开销——比如分配线程栈、切换调度上下文,这些成本对于“push_back一个int”这种超轻量任务来说,完全是本末倒置,测试规模越大,线程数量越多,这部分开销就越爆炸。

2. 全局锁把并行彻底变成串行

所有线程都在争抢同一个the_lock,每次只有一个线程能执行push_back,其余线程都处于等待锁的状态。这相当于把多线程逻辑硬生生变成了串行执行,还额外叠加了锁的等待、释放开销,比单线程直接跑慢是必然结果。

3. 任务粒度太小

单个push_back操作耗时微乎其微,并行带来的收益完全覆盖不了线程和锁的额外开销。并行技术更适合处理计算密集、粒度足够大的任务,小任务批量处理才能体现出并行的价值。


优化方案

针对上述问题,调整思路:按CPU核心数拆分任务,每个线程处理一批元素,用局部vector避免锁竞争,最后一次性合并结果。

优化后的代码示例:

#include <iostream>
#include <future>
#include <chrono>
#include <vector>
#include <mutex>
#include <thread>

// 假设你的Time类实现正确,此处省略具体代码
class Time{
    // ... 你的原有实现
};

#define timer() Time timer##__LINE__

// 批量处理任务:每个线程处理[start, end)区间的元素
static void batch_addition(std::vector<int>* result, int start, int end) {
    std::vector<int> local_vec;
    local_vec.reserve(end - start); // 预分配空间,避免频繁扩容
    for (int i = start; i < end; ++i) {
        local_vec.push_back(i + 2);
    }
    // 最后一次性合并到全局vector,仅需加一次锁
    static std::mutex the_lock;
    std::lock_guard<std::mutex> lock(the_lock);
    result->insert(result->end(), local_vec.begin(), local_vec.end());
}

int main()
{
    std::vector<int> array;
    int test_size = 10000;
    std::vector<std::future<void>> futures;

    // 获取CPU核心数,确定合理的线程数量
    int thread_count = std::thread::hardware_concurrency();
    if (thread_count == 0) thread_count = 4; // 兜底默认值

    // 拆分任务到各个线程
    {
        timer();
        int batch_size = test_size / thread_count;
        for (int i = 0; i < thread_count; ++i) {
            int start = i * batch_size;
            int end = (i == thread_count - 1) ? test_size : (i + 1) * batch_size;
            futures.push_back(std::async(std::launch::async, batch_addition, &array, start, end));
        }
        // 等待所有线程执行完成
        for (auto& f : futures) {
            f.get();
        }
    }

    // 单线程版本对比(预分配空间保证公平性)
    {
        timer();
        array.clear();
        array.reserve(test_size);
        for(int i = 0; i < test_size; ++i){
            array.push_back(i + 2);
        }
    }
    
    return 0;
}

优化效果说明

  • 线程数量控制在CPU核心数级别,避免大量线程创建的冗余开销。
  • 每个线程先处理局部vector,最后一次性合并,锁的开销降到最低。
  • 任务粒度变大后,并行的计算收益能覆盖线程开销,当测试规模足够大时,异步版本的性能会追上甚至超过单线程。
  • 给vector预分配reserve空间,避免扩容带来的额外开销,让单线程与异步版本的对比更公平。

内容的提问来源于stack exchange,提问作者Page Mike

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.03 21:29:51