如何高效统计超大文本文件中每行的出现次数?
优化大文件行重复次数统计的方案
你的代码运行缓慢核心是三个问题:
- 用
readlines()一次性把几百万行加载到内存,既占资源又拖慢处理速度 - 手动调用
yourResult.count(i)统计次数,时间复杂度是O(n²),数据量大时完全不可用 - 做了大量无意义操作(比如
split('\n')、sum合并列表),徒增额外开销
下面是几个逐层优化的实现方式:
1. 最简洁的Python优化版
直接用Counter逐行读取文件,无需把所有内容加载到内存:
from collections import Counter def count_lines(file_path): with open(file_path, 'r', encoding='utf-8') as f: # 用生成器逐行处理,自动迭代文件,内存占用极低 line_counter = Counter(line.rstrip('\n') for line in f) # 按出现次数降序排序 sorted_result = sorted(line_counter.items(), key=lambda x: x[1], reverse=True) return sorted_result # 使用示例 result = count_lines(r"C:\temp\large text_file.txt") for line, count in result: print(f"{count}: {line}")
这里的核心是用生成器表达式给Counter,不需要把所有行存入列表,统计效率为O(n),内存占用仅为单行道级别。
2. 超大规模文件适配的手动统计
如果文件大到连Counter的内存占用都有压力,可改用defaultdict手动计数,逻辑更灵活:
from collections import defaultdict def count_lines(file_path): line_counts = defaultdict(int) with open(file_path, 'r', encoding='utf-8') as f: for line in f: # 仅删除换行符(如需忽略首尾空白可换成strip()) cleaned_line = line.rstrip('\n') line_counts[cleaned_line] += 1 sorted_result = sorted(line_counts.items(), key=lambda x: x[1], reverse=True) return sorted_result
这个方案和Counter效率接近,但可以随时插入自定义逻辑(比如过滤空行、处理特殊字符)。
3. 比Python更快的命令行方案
如果你的系统是Linux/macOS,或Windows上有Git Bash/WSL,直接用原生命令行工具速度会远超Python:
# 先排序、再统计重复行、最后按次数降序输出 sort "C:\temp\large text_file.txt" | uniq -c | sort -nr
如需忽略大小写或首尾空白,可添加参数:
# 忽略大小写+忽略首尾空白 sort -f "C:\temp\large text_file.txt" | sed 's/^[[:space:]]*//;s/[[:space:]]*$//' | uniq -ic | sort -nr
原代码问题拆解
readlines():会把整个文件加载到内存,几百万行场景下极易触发内存溢出,换成for line in f逐行迭代才是正确姿势line.strip().split('\n'):line本身就是单行内容,split('\n')会生成单元素列表,完全多余;后续sum(yourResult, [])合并列表效率极低(每次合并都会创建新列表)dict((i, yourResult.count(i)) for i in yourResult):每个元素都要遍历整个列表统计,时间复杂度O(n²),几百万行场景下这个操作会慢到无法忍受,而Counter内部用哈希表统计,时间复杂度仅为O(n)
内容的提问来源于stack exchange,提问作者Mark K
相关产品推荐
相关产品推荐

