Python多文件文本清洗:合并处理与词频统计最优方案咨询
处理多TXT文件:先合并还是分别处理?
嗨 Sarah,我来帮你拆解这个问题,顺便把你的代码升级成支持多文件的版本~
效率对比:两种方案各有什么优劣势?
- 先合并文本再统一处理:
- 优点:代码逻辑更简单,不用维护多个词频字典的合并逻辑,对于中小型文件(总大小在几GB以内)来说,内存占用完全可控,执行效率也很高。
- 缺点:如果你的文件总大小特别大(比如几十GB),一次性加载所有文本到内存可能会撑爆内存,这时候就需要分批处理。
- 先分别处理每个文件再合并词频:
- 优点:内存占用更低,每次只处理单个文件,适合超大规模的文件集合。
- 缺点:代码复杂度更高,需要额外写字典合并的逻辑(比如把每个文件的Counter合并到总Counter里),而且多次文件IO的开销会比一次性合并略高一点。
最优选择:如果你的文件都是普通大小(日常文本分析场景大多是这样),优先选「先合并再处理」,代码简洁好维护,效率也足够。如果是超大规模文件,再考虑分批处理。
升级后的代码(先合并再处理版本)
我基于你的原代码做了优化,修复了重复导入的问题,改用with语句安全操作文件,同时支持批量读取指定目录下的所有.txt文件:
import re import string import csv from collections import Counter from glob import glob # 匹配指定目录下的所有txt文件,这里可以改成你的文件路径,比如 './data/*.txt' txt_files = glob('*.txt') # 当前目录下所有txt文件 # 合并所有文件的文本 combined_text = "" for file_path in txt_files: with open(file_path, 'rt', encoding='utf-8') as f: combined_text += f.read() + " " # 加空格避免两个文件末尾连在一起 # 文本清洗与词列表生成 # 分割单词(非字母数字的字符作为分隔符) words = re.split(r'\W+', combined_text) # 转小写 words = [word.lower() for word in words] # 去除标点符号 table = str.maketrans('', '', string.punctuation) stripped_words = [w.translate(table) for w in words] # 过滤空字符串(分割后可能产生空值) stripped_words = [word for word in stripped_words if word] # 统计词频 word_counts = Counter(stripped_words) # 输出到CSV文件 with open('combined_word_counts.csv', 'w', newline='', encoding='utf-8') as csv_file: writer = csv.writer(csv_file) writer.writerow(['Word', 'Count']) # 写入表头 for word, count in word_counts.items(): writer.writerow([word, count]) print(f"处理完成!共合并了{len(txt_files)}个txt文件,结果已保存到combined_word_counts.csv")
代码说明
- 用
glob('*.txt')批量获取所有txt文件路径,你可以修改路径匹配规则(比如'./text_files/*.txt'指定子目录)。 - 用
with语句操作文件,自动关闭文件,避免资源泄漏。 - 增加了过滤空字符串的步骤,避免统计到分割产生的空内容。
- 加入了表头,让CSV文件更规范。
- 如果你的文件编码不是utf-8,可以修改
open里的encoding参数(比如gbk)。
如果你需要处理超大规模文件(分批处理版本)
如果你的文件总大小太大,没法一次性加载,就用这个版本,逐个处理文件并合并词频:
import re import string import csv from collections import Counter from glob import glob txt_files = glob('*.txt') total_counts = Counter() for file_path in txt_files: with open(file_path, 'rt', encoding='utf-8') as f: text = f.read() # 单文件清洗 words = re.split(r'\W+', text) words = [word.lower() for word in words] table = str.maketrans('', '', string.punctuation) stripped_words = [w.translate(table) for w in words] stripped_words = [word for word in stripped_words if word] # 合并到总词频 total_counts.update(stripped_words) print(f"已处理文件:{file_path}") # 输出CSV with open('combined_word_counts.csv', 'w', newline='', encoding='utf-8') as csv_file: writer = csv.writer(csv_file) writer.writerow(['Word', 'Count']) for word, count in total_counts.items(): writer.writerow([word, count]) print(f"所有文件处理完成!结果已保存到combined_word_counts.csv")
这个版本每次只加载一个文件的内容到内存,处理完就释放,适合超大文件场景。
内容的提问来源于stack exchange,提问作者Sarah
相关产品推荐
相关产品推荐

