多文本文件停用词过滤脚本异常:输出均为None,求修复方案
问题
我有一个cleaned_texts文件夹,里面是多个文本文件(比如a.txt、b.txt),每个文件内容是分词后的单词列表,格式示例:['Rise', 'of', 'e-health', 'and', 'its', 'Germany', 'dollar'];另外有StopWords文件夹,包含多类停用词文件(如currency.txt、geographic.txt),每个文件存对应类别停用词。
我需要把StopWords里的所有停用词,从cleaned_texts的所有文本文件中过滤掉。目前已经遍历StopWords合并出停用词列表new_list,但脚本运行后,生成的new_texts文件夹里的文件全是None,得不到预期过滤结果(比如a.txt预期是['Rise', 'of', 'e-health', 'and', 'its'])。
附上我的Python脚本:
import glob import codecs import os #Cleaned texts os.getcwd() clean_texts_folder = os.path.join(os.getcwd(), 'cleaned_texts') clean_text_data = [] for root, folders, files in os.walk(clean_texts_folder): for file in files: path = os.path.join(root, file) with codecs.open(path, encoding='utf-8', errors='ignore') as info: clean_text_data.append(info.read()) #Stop Words stopwords_folder_path = "StopWords" stopwords_files = glob.glob(os.path.join(stopwords_folder_path, '*.txt')) for file in stopwords_files: with open(file, 'r') as w: stop_words = w.read() map_dict = {'|': ''} res = ''.join( idx if idx not in map_dict else map_dict[idx] for idx in stop_words) new_list = res.split() #new_list Output= ['SMITH', 'Surnames', 'from', '1990', 'Thailand', 'YEN', 'India', 'PESO', 'Japan', 'Canada'] #Trying to save the filtered texts folder_name = "new_texts" Path(folder).mkdir(parents=True, exist_ok=True) filtered_sentence = [] for index, word in enumerate(clean_text_data): if word not in new_list: #print(filtered_sentence.append(word)) file_path = Path(folder_name, f"{index}.txt") with pathlib.Path.open(file_path, "w", encoding="utf-8") as f: f.write(f"{filtered_sentence }")
问题分析
- 停用词列表未合并:循环处理停用词文件时,每次都会覆盖
new_list,最终仅保留最后一个文件的停用词,未合并所有文件内容。 - 未解析列表格式:读取
cleaned_texts文件时直接存入字符串,未将字符串形式的列表转为Python实际列表,导致判断的是整个字符串而非单个单词。 - 过滤逻辑错误:将每个文件的全部内容当作单个
word处理,且filtered_sentence未按文件重置,写入时未正确收集过滤后的单词。 - 模块与路径错误:使用
Path和pathlib相关方法但未导入模块;创建文件夹时用了未定义的folder变量,应为folder_name。 - 写入逻辑混乱:循环中错误地创建文件并写入空列表,导致内容为空或异常。
修正后的完整脚本
import glob import codecs import os import pathlib from ast import literal_eval # 读取并合并所有停用词 stopwords_folder_path = "StopWords" stopwords_files = glob.glob(os.path.join(stopwords_folder_path, '*.txt')) new_list = [] for file in stopwords_files: with open(file, 'r', encoding='utf-8') as w: stop_words = w.read() # 替换|并分割单词,合并到总停用词列表 res = stop_words.replace('|', '') new_list.extend(res.split()) # 转为集合提升查找效率 stop_words_set = set(new_list) # 处理cleaned_texts中的每个文件 clean_texts_folder = os.path.join(os.getcwd(), 'cleaned_texts') folder_name = "new_texts" # 创建目标文件夹 pathlib.Path(folder_name).mkdir(parents=True, exist_ok=True) for root, folders, files in os.walk(clean_texts_folder): for file in files: file_path = os.path.join(root, file) with codecs.open(file_path, encoding='utf-8', errors='ignore') as info: # 将字符串列表转为Python实际列表 text_content = info.read().strip() word_list = literal_eval(text_content) # 过滤停用词 filtered_words = [word for word in word_list if word not in stop_words_set] # 按原文件名写入新文件夹 output_file_path = os.path.join(folder_name, file) with open(output_file_path, "w", encoding="utf-8") as f: f.write(str(filtered_words))
关键修正点说明
- 合并停用词:用
extend()替代直接赋值,将所有停用词文件的内容合并到总列表;转为集合后,成员查找效率远高于列表。 - 解析列表字符串:通过
ast.literal_eval()将文件中的字符串列表转为可操作的Python列表,实现逐个单词过滤。 - 保留原文件名:处理后用原文件名保存,方便对应原文件。
- 修复模块与路径:导入
pathlib模块,修正文件夹创建的变量错误。 - 独立文件过滤:对每个文件单独处理过滤逻辑,确保每个文件的结果独立正确。
内容的提问来源于stack exchange,提问作者Abuchi
相关产品推荐
相关产品推荐

