You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用C++17及以后版本的std::transform_reduce()实现并行词频统计?

使用std::transform_reduce()统计单词出现次数(C++17及以上)

完整可编译实现代码

#include <iostream>
#include <string>
#include <sstream>
#include <vector>
#include <unordered_map>
#include <execution>
#include <numeric>

int main() {
    std::string text = "apple orange banana apple apple orange";
    std::istringstream iss(text);
    std::vector words(std::istream_iterator<std::string>(iss), {});

    // 使用transform_reduce并行统计单词次数
    auto wordCount = std::transform_reduce(
        std::execution::par,  // 启用并行执行策略
        words.begin(), words.end(),
        std::unordered_map<std::string, int>{},  // 归约初始值:空哈希表
        // 归约操作:合并两个哈希表,累加单词计数
        [](std::unordered_map<std::string, int> lhs, const std::unordered_map<std::string, int>& rhs) {
            for (const auto& [word, count] : rhs) {
                lhs[word] += count;
            }
            return lhs;
        },
        // 转换操作:将单个单词转为仅含该单词、计数为1的小型哈希表
        [](const std::string& word) {
            return std::unordered_map<std::string, int>{{word, 1}};
        }
    );

    // 输出统计结果
    for (const auto& [word, count] : wordCount) {
        std::cout << word << ": " << count << std::endl;
    }

    return 0;
}

核心逻辑说明

  • 并行执行策略:std::execution::par 触发并行处理,编译器会自动将任务拆分到多个线程执行,适配大数据集场景;若无需并行,可替换为std::execution::seq(串行执行)。
  • 转换Lambda:将每个单词转换为仅包含该单词、计数为1的局部哈希表,确保每个线程独立处理子任务,避免全局数据竞争。
  • 归约Lambda:负责合并多个局部哈希表,将子统计结果的计数累加到全局哈希表中,并行执行时系统会自动完成多线程结果的合并。
  • 依赖头文件:必须包含<execution>(提供并行策略)和<numeric>(transform_reduce的定义),否则无法编译。

关键注意事项

  • 并行场景下禁止直接修改全局哈希表,会引发数据竞争导致未定义行为;通过局部哈希表再合并的方式,才能安全利用并行能力。
  • 确保编译器支持C++17并行标准库(如GCC 9+、Clang 10+、MSVC 2019+)。

内容的提问来源于stack exchange,提问作者Will

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.20 09:45:11