如何统计字符频次并按频次排序?含文件读取与编码前置需求
字符读取与频次统计实现方案
1. 字符读取(含所有符号、空格)
直接按字节读取文件,保留所有字符(包括空格、标点)。如果需要统一大小写(如示例中把I转为i),读取后用tolower处理字符。注意不要用>>运算符读取,它会自动跳过空白字符,改用file.get(c)读取每一个字符。
2. 频次统计
用std::unordered_map<char, int>做统计效率最高,遍历每个读取到的字符,对应键的计数自增即可。统计完成后转成vector<pair<char, int>>,方便后续排序或堆操作。
核心代码片段
#include <fstream> #include <unordered_map> #include <vector> #include <algorithm> #include <cctype> int main() { std::ifstream input_file("your_input.txt"); std::unordered_map<char, int> count_map; char curr_char; // 读取并统计所有字符 while (input_file.get(curr_char)) { // 统一转小写(不需要区分大小写则删除此行) curr_char = std::tolower(static_cast<unsigned char>(curr_char)); count_map[curr_char]++; } // 转换为vector,便于排序 std::vector<std::pair<char, int>> count_vec(count_map.begin(), count_map.end()); // 按频次从小到大排序,频次相同则按字符顺序排列 std::sort(count_vec.begin(), count_vec.end(), [](const auto& a, const auto& b) { if (a.second != b.second) { return a.second < b.second; } return a.first < b.first; }); // 输出结果(空格单独标注) for (const auto& item : count_vec) { if (item.first == ' ') { std::cout << ": " << item.second << " "; } else { std::cout << item.first << ": " << item.second << " "; } } return 0; }
3. 用堆排序实现需求
如果必须用堆排序替代std::sort,可以用std::make_heap和std::sort_heap实现最小堆排序:
// 构建最小堆(比较器控制堆顶为频次最小的元素) std::make_heap(count_vec.begin(), count_vec.end(), [](const auto& a, const auto& b) { if (a.second != b.second) { return a.second > b.second; } return a.first > b.first; }); // 对堆进行排序,结果为从小到大排列 std::sort_heap(count_vec.begin(), count_vec.end(), [](const auto& a, const auto& b) { if (a.second != b.second) { return a.second > b.second; } return a.first > b.first; });
常见问题排查
- 文件读取失败:检查文件路径是否正确、是否有读取权限。
- 大小写未统一:忘记调用
tolower,导致大小写字符被分开统计。 - 空格未统计:误用
file >> curr_char,该运算符会跳过空白字符,必须用file.get(curr_char)。 - 排序逻辑错误:比较器写反,导致频次从大到小排列,注意最小堆的比较器需让频次小的元素优先在堆顶。
内容的提问来源于stack exchange,提问作者Ariani P.
相关产品推荐
相关产品推荐

