如何用Python批量处理多文本文件,提取停用词并分文件保存结果?
批量文本文件停用词提取与处理修复方案
需求与问题
需要批量处理文件夹中的100+文本文件,每个源文件对应生成两个输出文件:
Stop_Word_Consist_原文件名.txt:存储该文件中实际出现的停用词Stop_word_not_原文件名.txt:存储去除停用词后的文本
原代码能生成文件,但停用词文件内容不符合预期,未正确提取目标文件中的停用词,核心问题出在逻辑、执行顺序及路径处理上。
原代码核心问题分析
- 函数逻辑错误:
clean_text未使用传入参数,反而引用全局data变量,导致处理所有文件而非单个文件 - 循环对象错误:
for entry in data遍历的是DataFrame列名,不是单个文件路径 - 执行顺序颠倒:先尝试写入文件再调用
clean_text,此时all_reviews和stop_words未初始化 - 停用词提取错误:返回通用停用词表,而非当前文件中实际出现的停用词
- 路径处理不规范:未过滤非文本文件,输出文件未保存到预先创建的
Stopwords_folder
修正后的完整代码
import os import re import string import nltk from nltk.tokenize import word_tokenize from nltk.corpus import stopwords from nltk.stem.porter import PorterStemmer # 下载必要的NLTK数据(首次运行需要) nltk.download('punkt') nltk.download('stopwords') # 配置路径 source_dir = os.getcwd() # 源文件所在文件夹 output_dir = os.path.join(source_dir, 'Stopwords_folder') os.makedirs(output_dir, exist_ok=True) # 初始化全局停用词表(排除not)和词干提取器 global_stop_words = set(stopwords.words("english")) global_stop_words.discard("not") ps = PorterStemmer() def process_single_file(file_path): """处理单个文本文件,返回去停用词后的文本和该文件中出现的停用词集合""" # 读取文件内容 with open(file_path, 'r', encoding='utf-8', errors='ignore') as f: text = f.read() # 文本预处理 text = text.lower() # 移除URL url_pattern = re.compile('http[s]?://(?:[a-zA-Z]|[0-9]|[$-_@.&+]|[!*\(\),]|(?:%[0-9a-fA-F][0-9a-fA-F]))+') text = url_pattern.sub('', text) # 移除emoji emoji_pattern = re.compile("[" u"\U0001F600-\U0001FFFF" # 表情符号 u"\U0001F300-\U0001F5FF" # 符号与象形文字 u"\U0001F680-\U0001F6FF" # 交通与地图符号 u"\U0001F1E0-\U0001F1FF" # 国旗 u"\U00002702-\U000027B0" u"\U000024C2-\U0001F251" "]+", flags=re.UNICODE) text = emoji_pattern.sub(r'', text) # 缩写替换 contractions = { r"i'm": "i am", r"he's": "he is", r"she's": "she is", r"that's": "that is", r"what's": "what is", r"where's": "where is", r"\'ll": " will", r"\'ve": " have", r"\'re": " are", r"\'d": " would", r"won't": "will not", r"don't": "do not", r"did't": "did not", r"can't": "can not", r"it's": "it is", r"couldn't": "could not", r"have't": "have not" } for contraction, replacement in contractions.items(): text = re.sub(contraction, replacement, text) # 移除特殊符号 text = re.sub(r"[,.\"!@#$%^&*(){}?/;`~:<>=+-]", "", text) # 分词与过滤 tokens = word_tokenize(text) # 移除标点 table = str.maketrans('', '', string.punctuation) stripped_tokens = [w.translate(table) for w in tokens] # 只保留字母单词 alpha_tokens = [word for word in stripped_tokens if word.isalpha()] # 提取当前文件中出现的停用词 file_stop_words = set() filtered_words = [] for word in alpha_tokens: if word in global_stop_words: file_stop_words.add(word) else: filtered_words.append(ps.stem(word)) filtered_text = ' '.join(filtered_words) return filtered_text, file_stop_words # 批量处理所有txt文件 for filename in os.listdir(source_dir): if filename.endswith('.txt'): file_path = os.path.join(source_dir, filename) filtered_text, file_stop_words = process_single_file(file_path) # 生成输出文件名 no_stop_filename = f"Stop_word_not_{filename}" stop_word_filename = f"Stop_Word_Consist_{filename}" # 写入去停用词后的文件 with open(os.path.join(output_dir, no_stop_filename), 'w', encoding='utf-8') as f: f.write(filtered_text) # 写入停用词文件 with open(os.path.join(output_dir, stop_word_filename), 'w', encoding='utf-8') as f: f.write(' '.join(file_stop_words)) print("批量处理完成!")
关键修正点说明
- 函数重构:
process_single_file专注处理单个文件,逻辑独立清晰 - 停用词提取修正:实时收集当前文件中实际出现的停用词,而非返回全局表
- 循环逻辑优化:仅遍历
.txt后缀文件,避免处理无关文件 - 执行顺序调整:先处理文件得到结果,再写入输出文件
- 路径规范:使用
os.path.join处理路径,适配跨平台环境,输出文件统一存入指定文件夹 - 资源优化:提前初始化停用词表和词干提取器,避免重复创建资源
内容的提问来源于stack exchange,提问作者ANISH GAIKWAD
相关产品推荐
相关产品推荐

