You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python多文件文本清洗:合并处理与词频统计最优方案咨询

处理多TXT文件:先合并还是分别处理?

嗨 Sarah,我来帮你拆解这个问题,顺便把你的代码升级成支持多文件的版本~

效率对比:两种方案各有什么优劣势?

  • 先合并文本再统一处理:
    • 优点:代码逻辑更简单,不用维护多个词频字典的合并逻辑,对于中小型文件(总大小在几GB以内)来说,内存占用完全可控,执行效率也很高。
    • 缺点:如果你的文件总大小特别大(比如几十GB),一次性加载所有文本到内存可能会撑爆内存,这时候就需要分批处理。
  • 先分别处理每个文件再合并词频:
    • 优点:内存占用更低,每次只处理单个文件,适合超大规模的文件集合。
    • 缺点:代码复杂度更高,需要额外写字典合并的逻辑(比如把每个文件的Counter合并到总Counter里),而且多次文件IO的开销会比一次性合并略高一点。

最优选择:如果你的文件都是普通大小(日常文本分析场景大多是这样),优先选「先合并再处理」,代码简洁好维护,效率也足够。如果是超大规模文件,再考虑分批处理。

升级后的代码(先合并再处理版本)

我基于你的原代码做了优化,修复了重复导入的问题,改用with语句安全操作文件,同时支持批量读取指定目录下的所有.txt文件:

import re
import string
import csv
from collections import Counter
from glob import glob

# 匹配指定目录下的所有txt文件,这里可以改成你的文件路径,比如 './data/*.txt'
txt_files = glob('*.txt')  # 当前目录下所有txt文件

# 合并所有文件的文本
combined_text = ""
for file_path in txt_files:
    with open(file_path, 'rt', encoding='utf-8') as f:
        combined_text += f.read() + " "  # 加空格避免两个文件末尾连在一起

# 文本清洗与词列表生成
# 分割单词(非字母数字的字符作为分隔符)
words = re.split(r'\W+', combined_text)
# 转小写
words = [word.lower() for word in words]
# 去除标点符号
table = str.maketrans('', '', string.punctuation)
stripped_words = [w.translate(table) for w in words]
# 过滤空字符串(分割后可能产生空值)
stripped_words = [word for word in stripped_words if word]

# 统计词频
word_counts = Counter(stripped_words)

# 输出到CSV文件
with open('combined_word_counts.csv', 'w', newline='', encoding='utf-8') as csv_file:
    writer = csv.writer(csv_file)
    writer.writerow(['Word', 'Count'])  # 写入表头
    for word, count in word_counts.items():
        writer.writerow([word, count])

print(f"处理完成!共合并了{len(txt_files)}个txt文件,结果已保存到combined_word_counts.csv")

代码说明

  • 用glob('*.txt')批量获取所有txt文件路径,你可以修改路径匹配规则(比如'./text_files/*.txt'指定子目录)。
  • 用with语句操作文件,自动关闭文件,避免资源泄漏。
  • 增加了过滤空字符串的步骤,避免统计到分割产生的空内容。
  • 加入了表头,让CSV文件更规范。
  • 如果你的文件编码不是utf-8,可以修改open里的encoding参数(比如gbk)。

如果你需要处理超大规模文件(分批处理版本)

如果你的文件总大小太大,没法一次性加载,就用这个版本,逐个处理文件并合并词频:

import re
import string
import csv
from collections import Counter
from glob import glob

txt_files = glob('*.txt')
total_counts = Counter()

for file_path in txt_files:
    with open(file_path, 'rt', encoding='utf-8') as f:
        text = f.read()
        # 单文件清洗
        words = re.split(r'\W+', text)
        words = [word.lower() for word in words]
        table = str.maketrans('', '', string.punctuation)
        stripped_words = [w.translate(table) for w in words]
        stripped_words = [word for word in stripped_words if word]
        # 合并到总词频
        total_counts.update(stripped_words)
    print(f"已处理文件:{file_path}")

# 输出CSV
with open('combined_word_counts.csv', 'w', newline='', encoding='utf-8') as csv_file:
    writer = csv.writer(csv_file)
    writer.writerow(['Word', 'Count'])
    for word, count in total_counts.items():
        writer.writerow([word, count])

print(f"所有文件处理完成!结果已保存到combined_word_counts.csv")

这个版本每次只加载一个文件的内容到内存,处理完就释放,适合超大文件场景。

内容的提问来源于stack exchange,提问作者Sarah

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 07:04:04