You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

C++多线程优化遇性能差异:Boost版本及跨平台编译问题求助

C++多线程优化的两个性能问题解析

实现方案概述

我尝试通过多线程优化C++代码运行时间,实现了两种方案:

基于Boost.Asio的实现

#include <iostream>
#include <vector>
#include <chrono>
#include <atomic>
#include <thread>
#include <boost/asio.hpp>
#include <boost/bind/bind.hpp>

void Test::boost_worker_task() {
    char new_state[3][3];
    MyEngine::random_start_state(new_state);
    MyEngine::solve_game(new_state);
    ++games_solved_counter;
}

void Test::run(const unsigned int games_to_solve, const bool use_mul_thread) {
    MyEngine::init_rand();
    const auto start = std::chrono::high_resolution_clock::now();

    const unsigned int num_threads = use_mul_thread ? std::thread::hardware_concurrency() : 1;
    std::cout << "Using " << num_threads << " threads to solve " << games_to_solve << " games" << std::endl;

    boost::asio::io_service io_service;
    boost::asio::thread_pool pool(num_threads);

    for (unsigned int i = 0; i < games_to_solve; ++i) {
            io_service.post([] { return Test::boost_worker_task(); });
    }

    // Run and wait for all tasks to complete
    io_service.run();
    pool.join();

    const auto end = std::chrono::high_resolution_clock::now();
    const std::chrono::duration<double, std::milli> elapsed = end - start;

    std::cout << "Solved " << games_solved_counter << " games!" << std::endl;
    std::cout << "Elapsed time: " << elapsed.count() / 1000.0 << " seconds" << std::endl;
    std::cout << "Elapsed time: " << elapsed.count() << " milliseconds\n" << std::endl;
}

基于std::async的实现

#include <iostream>
#include <vector>
#include <thread>
#include <mutex>
#include <chrono>
#include <atomic>
#include <future>

std::atomic<int> games_solved(0);

void TestMulThread::worker_task(const unsigned int num_iterations, std::mutex& games_solved_mutex) {
    for (unsigned int i = 0; i < num_iterations; ++i) {
        char new_state[3][3];
        MyEngine::random_start_state(new_state);
        MyEngine::solve_game(new_state);

        ++games_solved;
    }
}
void TestMulThread::run(const unsigned int total_games_to_solve) {
    MyEngine::init_rand();
    const auto start_time = std::chrono::high_resolution_clock::now();
    const unsigned int num_threads = std::thread::hardware_concurrency();
    const unsigned int games_per_thread = total_games_to_solve / num_threads;
    const unsigned int remaining_games = total_games_to_solve % games_per_thread;
        std::cout << "Using " << num_threads << " threads to solve " << total_games_to_solve << " games" << std::endl;

    // Distribute the remaining games
    std::vector<unsigned int> games_for_each_thread(num_threads, games_per_thread);
    for (unsigned int i = 0; i < remaining_games; ++i) {
        games_for_each_thread[i]++;
    }

    std::vector<std::future<void>> futures;
    std::mutex games_solved_mutex;

    for (unsigned int i = 0; i < num_threads; ++i) {
        futures.push_back(std::async(std::launch::async, worker_task, games_for_each_thread[i], std::ref(games_solved_mutex)));
    }

    for (auto& future : futures) {
        future.get();
    }

    const auto end_time = std::chrono::high_resolution_clock::now();
    const auto elapsed_time = std::chrono::duration_cast<std::chrono::milliseconds>(end_time - start_time).count();
    std::cout << "Solved " << games_solved << " games!" << std::endl;
    std::cout << "Elapsed time: " << elapsed_time / 1000.0 << " seconds" << std::endl;
    std::cout << "Elapsed time: " << elapsed_time << " milliseconds\n" << std::endl;
}

遇到的问题及解析

问题1:Boost.Asio实现远慢于std::async甚至单线程

原因分析

  • 任务粒度与调度开销不匹配:给Boost.Asio提交了games_to_solve个独立小任务,每个任务仅执行一次游戏求解和原子计数。调度本身的开销远超任务执行时间,导致整体性能被拖垮。而std::async的实现是每个线程批量处理任务,调度开销被分摊到多个求解操作上,效率更高。
  • 线程池使用错误:代码中同时创建了io_service和thread_pool,但未将两者关联。默认情况下io_service.run()仅在当前线程执行任务,创建的thread_pool完全未被使用,所有任务串行执行,再加上调度的额外开销,自然比原生单线程循环还慢。

修复方案

  1. 调整任务粒度:将多个求解任务打包成大任务提交,减少调度次数。
  2. 正确使用线程池:直接用thread_pool的post方法提交批量任务,无需单独的io_service,示例修改:
// 替换原io_service和pool相关代码
boost::asio::thread_pool pool(num_threads);
unsigned int games_per_thread = games_to_solve / num_threads;
unsigned int remaining = games_to_solve % num_threads;

for (unsigned int i = 0; i < num_threads; ++i) {
    unsigned int count = games_per_thread + (i < remaining ? 1 : 0);
    boost::asio::post(pool, [this, count]() {
        for (unsigned int j = 0; j < count; ++j) {
            boost_worker_task();
        }
    });
}
pool.join();

问题2:std::async实现VS下快,Linux下g++编译后慢于单线程

原因分析

  • std::async调度行为差异:VS标准库会直接创建新线程执行任务,而g的libstdc可能存在额外调度开销或线程复用逻辑,导致任务执行效率下降。
  • 原子操作缓存开销:频繁的++games_solved会引发CPU缓存行频繁失效(伪共享)。VS编译器可能对原子操作做了合并或缓存友好优化,而g++实现更保守,缓存同步开销更大。
  • 线程调度策略差异:Linux默认线程调度可能导致线程频繁切换核心,降低缓存命中率;VS的调度更倾向于让线程固定在核心上运行,缓存效率更高。
  • 编译器优化细节不同:O2优化下,VS对循环展开、函数内联的优化程度可能高于g++,未充分挖掘代码性能。

修复方案

  1. 减少原子操作频率:每个线程用本地计数器统计,最后一次性更新全局原子变量,避免频繁缓存同步:
void TestMulThread::worker_task(const unsigned int num_iterations) {
    int local_count = 0;
    for (unsigned int i = 0; i < num_iterations; ++i) {
        char new_state[3][3];
        MyEngine::random_start_state(new_state);
        MyEngine::solve_game(new_state);
        local_count++;
    }
    games_solved += local_count; // 仅一次原子操作
}
  1. 调整编译参数:用更高优化等级-O3+针对当前CPU架构优化-march=native:
g++ -O3 -march=native -o test Test.cpp -std=c++20 -lpthread
  1. 用原生线程替代std::async:直接创建std::thread对象,避免std::async的调度不确定性:
std::vector<std::thread> threads;
for (unsigned int i = 0; i < num_threads; ++i) {
    threads.emplace_back(worker_task, games_for_each_thread[i]);
}
for (auto& t : threads) {
    t.join();
}

内容的提问来源于stack exchange,提问作者Jonas Örnfelt

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.02 13:34:55