You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

C++多线程文本统计代码无法并行运行的问题排查与优化

问题描述

我写了一段C++多线程代码,用来统计长文本文件metin.txt的单词数、句子数和段落数,每个统计功能对应一个独立线程,共3个线程。但运行后发现线程似乎不是同步并行执行,而是按顺序运行,找不到原因。我在Linux系统下运行这段代码,希望得到以下帮助:

  • 如何修改代码让所有线程真正同时运行?
  • 可以添加哪些内容进一步优化代码?
  • 如果当前代码不符合多线程逻辑,需要做哪些调整适配?
当前代码
#include <iostream>
#include <fstream>
#include <thread>
#include <mutex>
#include <cctype>
#include <chrono>

#define MAX_THREADS 3

std::string file_path = "metin.txt";
int total_words = 0;
int total_sentences = 0;
int total_paragraphs = 0;

std::mutex lock;

void count_words() {
    std::ifstream file(file_path);
    if (!file) {
        std::cout << "无法打开文件。" << std::endl;
        return;
    }

    int words = 0;
    char c;
    while ((c = file.get()) != EOF) {
        if (std::isspace(c)) {
            words++;
        }
    }

    lock.lock();
    total_words += words;
    lock.unlock();

    file.close();
}

void count_sentences() {
    std::ifstream file(file_path);
    if (!file) {
        std::cout << "无法打开文件。" << std::endl;
        return;
    }

    int sentences = 0;
    char c;
    while ((c = file.get()) != EOF) {
        if (c == '.') {
            sentences++;
        }
    }

    lock.lock();
    total_sentences += sentences;
    lock.unlock();

    file.close();
}

void count_paragraphs() {
    std::ifstream file(file_path);
    if (!file) {
        std::cout << "无法打开文件。" << std::endl;
        return;
    }

    int paragraphs = 0;
    char c;
    while ((c = file.get()) != EOF) {
        if (c == '\n') {
            paragraphs++;
        }
    }

    lock.lock();
    total_paragraphs += paragraphs;
    lock.unlock();

    file.close();
}

int main() {
    auto start_time = std::chrono::steady_clock::now(); // 开始时间

    std::thread threads[MAX_THREADS];

    int i;
    for (i = 0; i < MAX_THREADS; i++) {
        if (i == 0) {
            threads[i] = std::thread(count_words);
        }
        else if (i == 1) {
            threads[i] = std::thread(count_sentences);
        }
        else {
            threads[i] = std::thread(count_paragraphs);
        }
    }

    for (i = 0; i < MAX_THREADS; i++) {
        threads[i].join();
    }

    auto end_time = std::chrono::steady_clock::now(); // 结束时间
    auto elapsed_time = std::chrono::duration_cast<std::chrono::milliseconds>(end_time - start_time); // 耗时

    std::cout << "总单词数: " << total_words << std::endl;
    std::cout << "总句子数: " << total_sentences << std::endl;
    std::cout << "总段落数: " << total_paragraphs << std::endl;
    std::cout << "程序运行时间: " << elapsed_time.count() << " ms" << std::endl; // 输出耗时

    return 0;
}
问题分析与解决方案

1. 线程看似顺序执行的原因

  • 文件IO与缓存:每个线程都独立打开并读取整个文件,Linux系统的文件缓存机制会让第一个线程读取文件后,后续线程直接从内存缓存读取,速度极快,导致看起来像是顺序执行。
  • 线程调度特性:操作系统的线程调度是抢占式的,但如果任务执行时间极短,调度器可能还没来得及切换线程,任务就已完成,表现为顺序执行。

2. 让线程真正并行运行的核心调整

核心是避免重复读取文件,一次性将文件内容加载到内存,然后让三个线程共享内存数据进行统计,消除IO瓶颈,让线程真正并行处理计算任务。

3. 代码逻辑调整与优化

关键修正点

(1) 一次性加载文件到内存

在主线程中先读取整个文件内容到内存容器(比如std::string),然后传递给各个统计线程,避免每个线程重复打开文件。

(2) 修正统计逻辑的错误

  • 单词统计:当前按空格计数会把连续空格多次统计,正确逻辑是判断是否从非单词状态进入单词状态时计数。
  • 句子统计:仅统计.不全面,需包含!、?,同时避免缩写中的.(简单版可先处理常见结束符)。
  • 段落统计:当前按单个换行计数,正确逻辑是统计连续换行分隔的块,比如两个\n之间的内容算一个段落,开头有内容则默认第一个段落。

(3) 优化线程同步与变量管理

  • 使用std::lock_guard代替手动lock/unlock,确保异常情况下也能正确释放锁。
  • 避免全局变量,通过引用传递结果或者让线程返回统计值(用std::future),减少全局状态依赖。

调整后的代码示例

#include <iostream>
#include <fstream>
#include <thread>
#include <mutex>
#include <cctype>
#include <chrono>
#include <string>
#include <algorithm>

// 统计结果结构体,避免全局变量
struct Stats {
    int words = 0;
    int sentences = 0;
    int paragraphs = 0;
};

// 修正后的单词统计逻辑
void count_words(const std::string& content, Stats& stats, std::mutex& mtx) {
    int count = 0;
    bool in_word = false;

    for (char c : content) {
        if (std::isspace(c)) {
            in_word = false;
        } else if (!in_word) {
            in_word = true;
            count++;
        }
    }

    std::lock_guard<std::mutex> lock(mtx);
    stats.words = count;
}

// 修正后的句子统计逻辑
void count_sentences(const std::string& content, Stats& stats, std::mutex& mtx) {
    int count = 0;
    const char sentence_ends[] = {'.', '!', '?'};

    for (char c : content) {
        if (std::find(std::begin(sentence_ends), std::end(sentence_ends), c) != std::end(sentence_ends)) {
            count++;
        }
    }

    std::lock_guard<std::mutex> lock(mtx);
    stats.sentences = count;
}

// 修正后的段落统计逻辑
void count_paragraphs(const std::string& content, Stats& stats, std::mutex& mtx) {
    int count = 0;
    bool in_paragraph = false;

    for (char c : content) {
        if (c == '\n') {
            if (in_paragraph) {
                count++;
                in_paragraph = false;
            }
        } else if (!in_paragraph) {
            in_paragraph = true;
        }
    }
    // 最后一段如果有内容,补充计数
    if (in_paragraph) count++;

    std::lock_guard<std::mutex> lock(mtx);
    stats.paragraphs = count;
}

// 读取文件到内存
bool read_file(const std::string& path, std::string& content) {
    std::ifstream file(path, std::ios::binary);
    if (!file) {
        std::cerr << "无法打开文件: " << path << std::endl;
        return false;
    }

    // 读取全部内容
    content.assign((std::istreambuf_iterator<char>(file)), std::istreambuf_iterator<char>());
    return true;
}

int main() {
    auto start_time = std::chrono::steady_clock::now();

    std::string file_path = "metin.txt";
    std::string content;
    if (!read_file(file_path, content)) {
        return 1;
    }

    Stats stats;
    std::mutex mtx;

    // 创建线程,直接传递内存内容和统计结构体引用
    std::thread word_thread(count_words, std::cref(content), std::ref(stats), std::ref(mtx));
    std::thread sentence_thread(count_sentences, std::cref(content), std::ref(stats), std::ref(mtx));
    std::thread paragraph_thread(count_paragraphs, std::cref(content), std::ref(stats), std::ref(mtx));

    // 等待所有线程完成
    word_thread.join();
    sentence_thread.join();
    paragraph_thread.join();

    auto end_time = std::chrono::steady_clock::now();
    auto elapsed_time = std::chrono::duration_cast<std::chrono::milliseconds>(end_time - start_time);

    std::cout << "总单词数: " << stats.words << std::endl;
    std::cout << "总句子数: " << stats.sentences << std::endl;
    std::cout << "总段落数: " << stats.paragraphs << std::endl;
    std::cout << "程序运行时间: " << elapsed_time.count() << " ms" << std::endl;

    return 0;
}

4. 额外优化建议

  • 使用std::future替代共享结构体:让每个线程返回统计结果,主线程直接获取,无需互斥锁,进一步简化同步逻辑。
  • 大文件分块处理:如果文件极大,可将内存中的内容分块,让每个线程处理不同块,再汇总结果,提升并行效率。
  • 统计逻辑精细化:比如处理缩写中的.(如Mr.、Dr.)、省略号...等特殊情况,提升统计准确性。
  • 编译优化:在Linux下编译时添加-O2或-O3优化选项,提升代码运行效率。
  • 错误处理增强:添加更多文件读取、线程创建的错误检查,提升程序健壮性。

内容的提问来源于stack exchange,提问作者Fatih

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 18:32:34