You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

多文本文件停用词过滤脚本异常:输出均为None,求修复方案

问题

我有一个cleaned_texts文件夹,里面是多个文本文件(比如a.txt、b.txt),每个文件内容是分词后的单词列表,格式示例:['Rise', 'of', 'e-health', 'and', 'its', 'Germany', 'dollar'];另外有StopWords文件夹,包含多类停用词文件(如currency.txt、geographic.txt),每个文件存对应类别停用词。

我需要把StopWords里的所有停用词,从cleaned_texts的所有文本文件中过滤掉。目前已经遍历StopWords合并出停用词列表new_list,但脚本运行后,生成的new_texts文件夹里的文件全是None,得不到预期过滤结果(比如a.txt预期是['Rise', 'of', 'e-health', 'and', 'its'])。

附上我的Python脚本:

import glob
import codecs
import os

#Cleaned texts
os.getcwd()
clean_texts_folder =  os.path.join(os.getcwd(), 'cleaned_texts')

clean_text_data = []
for root, folders, files in os.walk(clean_texts_folder):
    for file in files:
        path = os.path.join(root, file)
        with codecs.open(path, encoding='utf-8', errors='ignore') as info:
            clean_text_data.append(info.read())


#Stop Words
stopwords_folder_path = "StopWords"
stopwords_files = glob.glob(os.path.join(stopwords_folder_path, '*.txt'))

for file in stopwords_files:
    with open(file, 'r') as w:
        stop_words = w.read()
        
        map_dict = {'|': ''}
        res = ''.join(
            idx if idx not in map_dict else map_dict[idx] for idx in stop_words)
        new_list = res.split()

#new_list Output= ['SMITH', 'Surnames', 'from', '1990', 'Thailand', 'YEN', 'India', 'PESO', 'Japan', 'Canada']


#Trying to save the filtered texts
folder_name = "new_texts"
Path(folder).mkdir(parents=True, exist_ok=True)
filtered_sentence = []
for index, word in enumerate(clean_text_data):
    if word not in new_list:
        #print(filtered_sentence.append(word))
        file_path = Path(folder_name, f"{index}.txt")
        with pathlib.Path.open(file_path, "w", encoding="utf-8") as f:
           f.write(f"{filtered_sentence }")

问题分析

  • 停用词列表未合并:循环处理停用词文件时,每次都会覆盖new_list,最终仅保留最后一个文件的停用词,未合并所有文件内容。
  • 未解析列表格式:读取cleaned_texts文件时直接存入字符串,未将字符串形式的列表转为Python实际列表,导致判断的是整个字符串而非单个单词。
  • 过滤逻辑错误:将每个文件的全部内容当作单个word处理,且filtered_sentence未按文件重置,写入时未正确收集过滤后的单词。
  • 模块与路径错误:使用Path和pathlib相关方法但未导入模块;创建文件夹时用了未定义的folder变量,应为folder_name。
  • 写入逻辑混乱:循环中错误地创建文件并写入空列表,导致内容为空或异常。

修正后的完整脚本

import glob
import codecs
import os
import pathlib
from ast import literal_eval

# 读取并合并所有停用词
stopwords_folder_path = "StopWords"
stopwords_files = glob.glob(os.path.join(stopwords_folder_path, '*.txt'))
new_list = []

for file in stopwords_files:
    with open(file, 'r', encoding='utf-8') as w:
        stop_words = w.read()
        # 替换|并分割单词,合并到总停用词列表
        res = stop_words.replace('|', '')
        new_list.extend(res.split())

# 转为集合提升查找效率
stop_words_set = set(new_list)

# 处理cleaned_texts中的每个文件
clean_texts_folder = os.path.join(os.getcwd(), 'cleaned_texts')
folder_name = "new_texts"
# 创建目标文件夹
pathlib.Path(folder_name).mkdir(parents=True, exist_ok=True)

for root, folders, files in os.walk(clean_texts_folder):
    for file in files:
        file_path = os.path.join(root, file)
        with codecs.open(file_path, encoding='utf-8', errors='ignore') as info:
            # 将字符串列表转为Python实际列表
            text_content = info.read().strip()
            word_list = literal_eval(text_content)
            
            # 过滤停用词
            filtered_words = [word for word in word_list if word not in stop_words_set]
            
            # 按原文件名写入新文件夹
            output_file_path = os.path.join(folder_name, file)
            with open(output_file_path, "w", encoding="utf-8") as f:
                f.write(str(filtered_words))

关键修正点说明

  • 合并停用词:用extend()替代直接赋值,将所有停用词文件的内容合并到总列表;转为集合后,成员查找效率远高于列表。
  • 解析列表字符串:通过ast.literal_eval()将文件中的字符串列表转为可操作的Python列表,实现逐个单词过滤。
  • 保留原文件名:处理后用原文件名保存,方便对应原文件。
  • 修复模块与路径:导入pathlib模块,修正文件夹创建的变量错误。
  • 独立文件过滤:对每个文件单独处理过滤逻辑,确保每个文件的结果独立正确。

内容的提问来源于stack exchange,提问作者Abuchi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.23 14:08:15