C++词频统计单词重复显示及文件误提示问题排查
C++单词统计程序的两个问题修复:大小写重复与错误提示误报
问题1:大小写单词重复统计(如the与The分别计数)
原因
C++中std::string的默认比较逻辑区分大小写,"the"和"The"会被当作两个独立的键存入map,导致重复统计。
解决方法
清理掉单词中的非字母字符后,将所有字符统一转换为小写(或大写),再存入map。需要引入<cctype>头文件来使用tolower函数。
修改代码片段:
// 清理非字母字符后添加: for (char& c : word) { c = tolower(static_cast<unsigned char>(c)); }
问题2:文件正常读取却打印"could not open file"提示
原因
代码仅在文件未打开时打印错误提示,但未终止程序。如果输出文件打开失败但输入文件正常,程序会继续执行统计逻辑,导致用户看到错误提示却有正常输出的矛盾情况。
解决方法
打印错误提示后立即终止程序,避免后续代码执行。同时优化提示信息,明确指出是输入还是输出文件打开失败。
修改代码片段:
if(!fs.is_open() || !output.is_open()){ if (!fs.is_open()) { cout << "could not open input file" << endl; } if (!output.is_open()) { cout << "could not open output file" << endl; } return 1; // 终止程序 }
另外,原代码中写入输出文件的循环错误使用cout而非output,需修正:
// 原错误代码: for(int i = 0; i < 30; i++){ cout << v[i].second << " : " << v[i].first << " times" << endl; } // 修改为: for(int i = 0; i < 30; i++){ output << v[i].second << " : " << v[i].first << " times" << endl; }
完整修改后的代码
#include <iostream> #include <vector> #include <map> #include <iterator> #include <fstream> #include <cctype> #include <algorithm> using namespace std; int main(){ fstream fs, output; fs.open("/Users/brah79/Downloads/skola/c++/inlämningsuppgifter/labb4/L4_wc/hitchhikersguide.txt"); output.open("/Users/brah79/Downloads/skola/c++/inlämningsuppgifter/labb4/labb4/output.txt"); if(!fs.is_open() || !output.is_open()){ if (!fs.is_open()) { cout << "could not open input file" << endl; } if (!output.is_open()) { cout << "could not open output file" << endl; } return 1; } map<string, int> mp; string word; while(fs >> word){ for(int i = 0; i < word.length(); i++){ if(!isalpha(static_cast<unsigned char>(word[i]))){ word.erase(i--, 1); } } if(word.empty()){ continue; } for (char& c : word) { c = tolower(static_cast<unsigned char>(c)); } mp[word]++; } vector<pair<int, string>> v; v.reserve(mp.size()); for (const auto& p : mp){ v.emplace_back(p.second, p.first); } sort(v.rbegin(), v.rend()); cout << "These are the 30 most frequent words: " << endl; for(int i = 0; i < 30 && i < v.size(); i++){ cout << v[i].second << " : " << v[i].first << " times" << endl; } output << "These are the 30 most frequent words: " << endl; for(int i = 0; i < 30 && i < v.size(); i++){ output << v[i].second << " : " << v[i].first << " times" << endl; } return 0; }
修复后的输出效果
重复的大小写单词会被合并统计,例如"the"的计数会变为2230+307=2537次,不再出现"The"单独条目;同时只有当文件确实打不开时才会打印错误提示并退出,不会出现矛盾输出。
内容的提问来源于stack exchange,提问作者sam
相关产品推荐
相关产品推荐

