You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python主题建模中文本预处理停用词未彻底清除的问题求助

Python主题建模中文本预处理停用词未彻底清除的问题求助

嘿,我最近在搞Python主题建模,为了把停用词清干净,特意把自己整理的停用词表、GitHub上找的一个停用词表,还有NLTK自带的停用词表合并在一起用了。结果跑完模型看主题的时候傻了眼——那些明明在停用词表里的词,比如"the""and"之类的,居然还好好地待在主题里,根本没被删掉!

我把预处理用的代码贴在下面了,代码运行起来完全没报错,但就是达不到效果,还给大家看个出问题的示例主题。我翻来覆去查了好几遍都没找到原因,有没有大佬能帮我诊断诊断?

我的预处理代码

import pandas as pd
import nltk
import spacy
from nltk.corpus import stopwords
from nltk.tokenize import word_tokenize
import requests

nltk.download('punkt')
nltk.download('stopwords')
nlp = spacy.load('en_core_web_sm')

# 加载GitHub停用词表
url = "https://gist.githubusercontent.com/rg089/35e00abf8941d72d419224cfd5b5925d/raw/12d899b70156fd0041fa9778d657330b024b959c/stopwords.txt"
github_stopwords = set(requests.get(url).text.splitlines())

# 自定义停用词(这里是我自己加的一些词)
custom_stopwords = {"your", "custom", "stopwords"}
# NLTK停用词
nltk_stopwords = set(stopwords.words('english'))
# 合并所有停用词表
all_stopwords = custom_stopwords.union(nltk_stopwords).union(github_stopwords)

# 预处理函数
def preprocess_text(text):
    # 分词
    tokens = word_tokenize(text)
    # 过滤停用词
    filtered_tokens = [token for token in tokens if token not in all_stopwords]
    return " ".join(filtered_tokens)

# 假设我用这个函数处理数据
# df['cleaned_text'] = df['text'].apply(preprocess_text)

出问题的示例主题

示例主题关键词:["the", "and", "to", "customer", "service"]
(这里"the""and""to"都是停用词,却没被过滤掉)

我自己琢磨的几个可能原因,也不确定对不对:

  1. 会不会是大小写没统一?比如停用词表里都是小写,但文本里有大写的词,匹配不上?
  2. 是不是GitHub的停用词表没正确加载?比如我刚才代码里有没有写错加载方式?
  3. 分词的问题?比如NLTK的word_tokenize把词拆成了奇怪的样子,和停用词不匹配?
  4. 是不是我预处理的时候没先去掉标点?比如文本里的"the,"带逗号,就和停用词"the"对不上?

有没有大佬能给我指条明路?

备注:内容来源于stack exchange,提问作者deniz

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.15 12:49:32