You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何高效统计超大文本文件中每行的出现次数?

优化大文件行重复次数统计的方案

你的代码运行缓慢核心是三个问题:

  • 用readlines()一次性把几百万行加载到内存,既占资源又拖慢处理速度
  • 手动调用yourResult.count(i)统计次数,时间复杂度是O(n²),数据量大时完全不可用
  • 做了大量无意义操作(比如split('\n')、sum合并列表),徒增额外开销

下面是几个逐层优化的实现方式:

1. 最简洁的Python优化版

直接用Counter逐行读取文件,无需把所有内容加载到内存:

from collections import Counter

def count_lines(file_path):
    with open(file_path, 'r', encoding='utf-8') as f:
        # 用生成器逐行处理,自动迭代文件,内存占用极低
        line_counter = Counter(line.rstrip('\n') for line in f)
    # 按出现次数降序排序
    sorted_result = sorted(line_counter.items(), key=lambda x: x[1], reverse=True)
    return sorted_result

# 使用示例
result = count_lines(r"C:\temp\large text_file.txt")
for line, count in result:
    print(f"{count}: {line}")

这里的核心是用生成器表达式给Counter,不需要把所有行存入列表,统计效率为O(n),内存占用仅为单行道级别。

2. 超大规模文件适配的手动统计

如果文件大到连Counter的内存占用都有压力,可改用defaultdict手动计数,逻辑更灵活:

from collections import defaultdict

def count_lines(file_path):
    line_counts = defaultdict(int)
    with open(file_path, 'r', encoding='utf-8') as f:
        for line in f:
            # 仅删除换行符(如需忽略首尾空白可换成strip())
            cleaned_line = line.rstrip('\n')
            line_counts[cleaned_line] += 1
    sorted_result = sorted(line_counts.items(), key=lambda x: x[1], reverse=True)
    return sorted_result

这个方案和Counter效率接近,但可以随时插入自定义逻辑(比如过滤空行、处理特殊字符)。

3. 比Python更快的命令行方案

如果你的系统是Linux/macOS,或Windows上有Git Bash/WSL,直接用原生命令行工具速度会远超Python:

# 先排序、再统计重复行、最后按次数降序输出
sort "C:\temp\large text_file.txt" | uniq -c | sort -nr

如需忽略大小写或首尾空白,可添加参数:

# 忽略大小写+忽略首尾空白
sort -f "C:\temp\large text_file.txt" | sed 's/^[[:space:]]*//;s/[[:space:]]*$//' | uniq -ic | sort -nr

原代码问题拆解

  • readlines():会把整个文件加载到内存,几百万行场景下极易触发内存溢出,换成for line in f逐行迭代才是正确姿势
  • line.strip().split('\n'):line本身就是单行内容,split('\n')会生成单元素列表,完全多余;后续sum(yourResult, [])合并列表效率极低(每次合并都会创建新列表)
  • dict((i, yourResult.count(i)) for i in yourResult):每个元素都要遍历整个列表统计,时间复杂度O(n²),几百万行场景下这个操作会慢到无法忍受,而Counter内部用哈希表统计,时间复杂度仅为O(n)

内容的提问来源于stack exchange,提问作者Mark K

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 22:55:29