C++多线程文本统计代码无法并行运行的问题排查与优化
问题描述
我写了一段C++多线程代码,用来统计长文本文件metin.txt的单词数、句子数和段落数,每个统计功能对应一个独立线程,共3个线程。但运行后发现线程似乎不是同步并行执行,而是按顺序运行,找不到原因。我在Linux系统下运行这段代码,希望得到以下帮助:
- 如何修改代码让所有线程真正同时运行?
- 可以添加哪些内容进一步优化代码?
- 如果当前代码不符合多线程逻辑,需要做哪些调整适配?
当前代码
#include <iostream> #include <fstream> #include <thread> #include <mutex> #include <cctype> #include <chrono> #define MAX_THREADS 3 std::string file_path = "metin.txt"; int total_words = 0; int total_sentences = 0; int total_paragraphs = 0; std::mutex lock; void count_words() { std::ifstream file(file_path); if (!file) { std::cout << "无法打开文件。" << std::endl; return; } int words = 0; char c; while ((c = file.get()) != EOF) { if (std::isspace(c)) { words++; } } lock.lock(); total_words += words; lock.unlock(); file.close(); } void count_sentences() { std::ifstream file(file_path); if (!file) { std::cout << "无法打开文件。" << std::endl; return; } int sentences = 0; char c; while ((c = file.get()) != EOF) { if (c == '.') { sentences++; } } lock.lock(); total_sentences += sentences; lock.unlock(); file.close(); } void count_paragraphs() { std::ifstream file(file_path); if (!file) { std::cout << "无法打开文件。" << std::endl; return; } int paragraphs = 0; char c; while ((c = file.get()) != EOF) { if (c == '\n') { paragraphs++; } } lock.lock(); total_paragraphs += paragraphs; lock.unlock(); file.close(); } int main() { auto start_time = std::chrono::steady_clock::now(); // 开始时间 std::thread threads[MAX_THREADS]; int i; for (i = 0; i < MAX_THREADS; i++) { if (i == 0) { threads[i] = std::thread(count_words); } else if (i == 1) { threads[i] = std::thread(count_sentences); } else { threads[i] = std::thread(count_paragraphs); } } for (i = 0; i < MAX_THREADS; i++) { threads[i].join(); } auto end_time = std::chrono::steady_clock::now(); // 结束时间 auto elapsed_time = std::chrono::duration_cast<std::chrono::milliseconds>(end_time - start_time); // 耗时 std::cout << "总单词数: " << total_words << std::endl; std::cout << "总句子数: " << total_sentences << std::endl; std::cout << "总段落数: " << total_paragraphs << std::endl; std::cout << "程序运行时间: " << elapsed_time.count() << " ms" << std::endl; // 输出耗时 return 0; }
问题分析与解决方案
1. 线程看似顺序执行的原因
- 文件IO与缓存:每个线程都独立打开并读取整个文件,Linux系统的文件缓存机制会让第一个线程读取文件后,后续线程直接从内存缓存读取,速度极快,导致看起来像是顺序执行。
- 线程调度特性:操作系统的线程调度是抢占式的,但如果任务执行时间极短,调度器可能还没来得及切换线程,任务就已完成,表现为顺序执行。
2. 让线程真正并行运行的核心调整
核心是避免重复读取文件,一次性将文件内容加载到内存,然后让三个线程共享内存数据进行统计,消除IO瓶颈,让线程真正并行处理计算任务。
3. 代码逻辑调整与优化
关键修正点
(1) 一次性加载文件到内存
在主线程中先读取整个文件内容到内存容器(比如std::string),然后传递给各个统计线程,避免每个线程重复打开文件。
(2) 修正统计逻辑的错误
- 单词统计:当前按空格计数会把连续空格多次统计,正确逻辑是判断是否从非单词状态进入单词状态时计数。
- 句子统计:仅统计
.不全面,需包含!、?,同时避免缩写中的.(简单版可先处理常见结束符)。 - 段落统计:当前按单个换行计数,正确逻辑是统计连续换行分隔的块,比如两个
\n之间的内容算一个段落,开头有内容则默认第一个段落。
(3) 优化线程同步与变量管理
- 使用
std::lock_guard代替手动lock/unlock,确保异常情况下也能正确释放锁。 - 避免全局变量,通过引用传递结果或者让线程返回统计值(用
std::future),减少全局状态依赖。
调整后的代码示例
#include <iostream> #include <fstream> #include <thread> #include <mutex> #include <cctype> #include <chrono> #include <string> #include <algorithm> // 统计结果结构体,避免全局变量 struct Stats { int words = 0; int sentences = 0; int paragraphs = 0; }; // 修正后的单词统计逻辑 void count_words(const std::string& content, Stats& stats, std::mutex& mtx) { int count = 0; bool in_word = false; for (char c : content) { if (std::isspace(c)) { in_word = false; } else if (!in_word) { in_word = true; count++; } } std::lock_guard<std::mutex> lock(mtx); stats.words = count; } // 修正后的句子统计逻辑 void count_sentences(const std::string& content, Stats& stats, std::mutex& mtx) { int count = 0; const char sentence_ends[] = {'.', '!', '?'}; for (char c : content) { if (std::find(std::begin(sentence_ends), std::end(sentence_ends), c) != std::end(sentence_ends)) { count++; } } std::lock_guard<std::mutex> lock(mtx); stats.sentences = count; } // 修正后的段落统计逻辑 void count_paragraphs(const std::string& content, Stats& stats, std::mutex& mtx) { int count = 0; bool in_paragraph = false; for (char c : content) { if (c == '\n') { if (in_paragraph) { count++; in_paragraph = false; } } else if (!in_paragraph) { in_paragraph = true; } } // 最后一段如果有内容,补充计数 if (in_paragraph) count++; std::lock_guard<std::mutex> lock(mtx); stats.paragraphs = count; } // 读取文件到内存 bool read_file(const std::string& path, std::string& content) { std::ifstream file(path, std::ios::binary); if (!file) { std::cerr << "无法打开文件: " << path << std::endl; return false; } // 读取全部内容 content.assign((std::istreambuf_iterator<char>(file)), std::istreambuf_iterator<char>()); return true; } int main() { auto start_time = std::chrono::steady_clock::now(); std::string file_path = "metin.txt"; std::string content; if (!read_file(file_path, content)) { return 1; } Stats stats; std::mutex mtx; // 创建线程,直接传递内存内容和统计结构体引用 std::thread word_thread(count_words, std::cref(content), std::ref(stats), std::ref(mtx)); std::thread sentence_thread(count_sentences, std::cref(content), std::ref(stats), std::ref(mtx)); std::thread paragraph_thread(count_paragraphs, std::cref(content), std::ref(stats), std::ref(mtx)); // 等待所有线程完成 word_thread.join(); sentence_thread.join(); paragraph_thread.join(); auto end_time = std::chrono::steady_clock::now(); auto elapsed_time = std::chrono::duration_cast<std::chrono::milliseconds>(end_time - start_time); std::cout << "总单词数: " << stats.words << std::endl; std::cout << "总句子数: " << stats.sentences << std::endl; std::cout << "总段落数: " << stats.paragraphs << std::endl; std::cout << "程序运行时间: " << elapsed_time.count() << " ms" << std::endl; return 0; }
4. 额外优化建议
- 使用
std::future替代共享结构体:让每个线程返回统计结果,主线程直接获取,无需互斥锁,进一步简化同步逻辑。 - 大文件分块处理:如果文件极大,可将内存中的内容分块,让每个线程处理不同块,再汇总结果,提升并行效率。
- 统计逻辑精细化:比如处理缩写中的
.(如Mr.、Dr.)、省略号...等特殊情况,提升统计准确性。 - 编译优化:在Linux下编译时添加
-O2或-O3优化选项,提升代码运行效率。 - 错误处理增强:添加更多文件读取、线程创建的错误检查,提升程序健壮性。
内容的提问来源于stack exchange,提问作者Fatih
相关产品推荐
相关产品推荐

