Python主题建模中文本预处理停用词未彻底清除的问题求助
Python主题建模中文本预处理停用词未彻底清除的问题求助
嘿,我最近在搞Python主题建模,为了把停用词清干净,特意把自己整理的停用词表、GitHub上找的一个停用词表,还有NLTK自带的停用词表合并在一起用了。结果跑完模型看主题的时候傻了眼——那些明明在停用词表里的词,比如"the""and"之类的,居然还好好地待在主题里,根本没被删掉!
我把预处理用的代码贴在下面了,代码运行起来完全没报错,但就是达不到效果,还给大家看个出问题的示例主题。我翻来覆去查了好几遍都没找到原因,有没有大佬能帮我诊断诊断?
我的预处理代码
import pandas as pd import nltk import spacy from nltk.corpus import stopwords from nltk.tokenize import word_tokenize import requests nltk.download('punkt') nltk.download('stopwords') nlp = spacy.load('en_core_web_sm') # 加载GitHub停用词表 url = "https://gist.githubusercontent.com/rg089/35e00abf8941d72d419224cfd5b5925d/raw/12d899b70156fd0041fa9778d657330b024b959c/stopwords.txt" github_stopwords = set(requests.get(url).text.splitlines()) # 自定义停用词(这里是我自己加的一些词) custom_stopwords = {"your", "custom", "stopwords"} # NLTK停用词 nltk_stopwords = set(stopwords.words('english')) # 合并所有停用词表 all_stopwords = custom_stopwords.union(nltk_stopwords).union(github_stopwords) # 预处理函数 def preprocess_text(text): # 分词 tokens = word_tokenize(text) # 过滤停用词 filtered_tokens = [token for token in tokens if token not in all_stopwords] return " ".join(filtered_tokens) # 假设我用这个函数处理数据 # df['cleaned_text'] = df['text'].apply(preprocess_text)
出问题的示例主题
示例主题关键词:["the", "and", "to", "customer", "service"]
(这里"the""and""to"都是停用词,却没被过滤掉)
我自己琢磨的几个可能原因,也不确定对不对:
- 会不会是大小写没统一?比如停用词表里都是小写,但文本里有大写的词,匹配不上?
- 是不是GitHub的停用词表没正确加载?比如我刚才代码里有没有写错加载方式?
- 分词的问题?比如NLTK的word_tokenize把词拆成了奇怪的样子,和停用词不匹配?
- 是不是我预处理的时候没先去掉标点?比如文本里的"the,"带逗号,就和停用词"the"对不上?
有没有大佬能给我指条明路?
备注:内容来源于stack exchange,提问作者deniz
相关产品推荐
相关产品推荐

