You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python批量处理多文本文件,提取停用词并分文件保存结果?

批量文本文件停用词提取与处理修复方案

需求与问题

需要批量处理文件夹中的100+文本文件,每个源文件对应生成两个输出文件:

  • Stop_Word_Consist_原文件名.txt:存储该文件中实际出现的停用词
  • Stop_word_not_原文件名.txt:存储去除停用词后的文本

原代码能生成文件,但停用词文件内容不符合预期,未正确提取目标文件中的停用词,核心问题出在逻辑、执行顺序及路径处理上。

原代码核心问题分析

  • 函数逻辑错误:clean_text未使用传入参数,反而引用全局data变量,导致处理所有文件而非单个文件
  • 循环对象错误:for entry in data遍历的是DataFrame列名,不是单个文件路径
  • 执行顺序颠倒:先尝试写入文件再调用clean_text,此时all_reviews和stop_words未初始化
  • 停用词提取错误:返回通用停用词表,而非当前文件中实际出现的停用词
  • 路径处理不规范:未过滤非文本文件,输出文件未保存到预先创建的Stopwords_folder

修正后的完整代码

import os
import re
import string
import nltk
from nltk.tokenize import word_tokenize
from nltk.corpus import stopwords
from nltk.stem.porter import PorterStemmer

# 下载必要的NLTK数据(首次运行需要)
nltk.download('punkt')
nltk.download('stopwords')

# 配置路径
source_dir = os.getcwd()  # 源文件所在文件夹
output_dir = os.path.join(source_dir, 'Stopwords_folder')
os.makedirs(output_dir, exist_ok=True)

# 初始化全局停用词表(排除not)和词干提取器
global_stop_words = set(stopwords.words("english"))
global_stop_words.discard("not")
ps = PorterStemmer()

def process_single_file(file_path):
    """处理单个文本文件,返回去停用词后的文本和该文件中出现的停用词集合"""
    # 读取文件内容
    with open(file_path, 'r', encoding='utf-8', errors='ignore') as f:
        text = f.read()
    
    # 文本预处理
    text = text.lower()
    # 移除URL
    url_pattern = re.compile('http[s]?://(?:[a-zA-Z]|[0-9]|[$-_@.&+]|[!*\(\),]|(?:%[0-9a-fA-F][0-9a-fA-F]))+')
    text = url_pattern.sub('', text)
    # 移除emoji
    emoji_pattern = re.compile("["
                               u"\U0001F600-\U0001FFFF"  # 表情符号
                               u"\U0001F300-\U0001F5FF"  # 符号与象形文字
                               u"\U0001F680-\U0001F6FF"  # 交通与地图符号
                               u"\U0001F1E0-\U0001F1FF"  # 国旗
                               u"\U00002702-\U000027B0"
                               u"\U000024C2-\U0001F251"
                               "]+", flags=re.UNICODE)
    text = emoji_pattern.sub(r'', text)
    # 缩写替换
    contractions = {
        r"i'm": "i am", r"he's": "he is", r"she's": "she is", r"that's": "that is",
        r"what's": "what is", r"where's": "where is", r"\'ll": " will", r"\'ve": " have",
        r"\'re": " are", r"\'d": " would", r"won't": "will not", r"don't": "do not",
        r"did't": "did not", r"can't": "can not", r"it's": "it is", r"couldn't": "could not",
        r"have't": "have not"
    }
    for contraction, replacement in contractions.items():
        text = re.sub(contraction, replacement, text)
    # 移除特殊符号
    text = re.sub(r"[,.\"!@#$%^&*(){}?/;`~:<>=+-]", "", text)
    
    # 分词与过滤
    tokens = word_tokenize(text)
    # 移除标点
    table = str.maketrans('', '', string.punctuation)
    stripped_tokens = [w.translate(table) for w in tokens]
    # 只保留字母单词
    alpha_tokens = [word for word in stripped_tokens if word.isalpha()]
    
    # 提取当前文件中出现的停用词
    file_stop_words = set()
    filtered_words = []
    for word in alpha_tokens:
        if word in global_stop_words:
            file_stop_words.add(word)
        else:
            filtered_words.append(ps.stem(word))
    
    filtered_text = ' '.join(filtered_words)
    return filtered_text, file_stop_words

# 批量处理所有txt文件
for filename in os.listdir(source_dir):
    if filename.endswith('.txt'):
        file_path = os.path.join(source_dir, filename)
        filtered_text, file_stop_words = process_single_file(file_path)
        
        # 生成输出文件名
        no_stop_filename = f"Stop_word_not_{filename}"
        stop_word_filename = f"Stop_Word_Consist_{filename}"
        
        # 写入去停用词后的文件
        with open(os.path.join(output_dir, no_stop_filename), 'w', encoding='utf-8') as f:
            f.write(filtered_text)
        
        # 写入停用词文件
        with open(os.path.join(output_dir, stop_word_filename), 'w', encoding='utf-8') as f:
            f.write(' '.join(file_stop_words))

print("批量处理完成!")

关键修正点说明

  • 函数重构:process_single_file专注处理单个文件,逻辑独立清晰
  • 停用词提取修正:实时收集当前文件中实际出现的停用词,而非返回全局表
  • 循环逻辑优化:仅遍历.txt后缀文件,避免处理无关文件
  • 执行顺序调整:先处理文件得到结果,再写入输出文件
  • 路径规范:使用os.path.join处理路径,适配跨平台环境,输出文件统一存入指定文件夹
  • 资源优化:提前初始化停用词表和词干提取器,避免重复创建资源

内容的提问来源于stack exchange,提问作者ANISH GAIKWAD

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 10:50:28