You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在pandas DataFrame列处理中应用自定义扩展停用词表

修改方法

你已经完成了自定义停用词和nltk原生停用词的集合合并,只需要把停用词过滤逻辑里引用的停用词变量,从原来的stop_words替换为合并后的new_stopwords_list即可,原有分词过滤的逻辑完全不需要调整。

完整可运行代码

from nltk.corpus import stopwords
import pandas as pd

# 加载nltk原生英文停用词表
stop_words = set(stopwords.words('english'))
# 自定义补充停用词列表
new_stopwords = ['satisfying', 'satisfy', 'satisfied', 'clemson', 'university', 'institution', 'disappointing', 'disappoint', 'disappointed', 'experience', 'would', 'should']
# 合并生成扩展停用词集合
new_stopwords_list = stop_words.union(new_stopwords)

# 过滤停用词,生成新列
df['stopwords_removed'] = df['no_punc'].apply(lambda x: [word for word in x if word not in new_stopwords_list])

df.head()

注意事项

  • 你的前置预处理已经把no_punc列的文本完成了分词、全小写转换,当前自定义停用词列表内的词均为小写格式,和待过滤文本格式完全匹配,不会出现漏匹配的问题。
  • 如果后续需要新增停用词,直接将对应小写格式的词汇加入new_stopwords列表即可,不需要修改后续过滤逻辑。
  • 如果首次运行nltk相关代码报错,可提前执行nltk.download('stopwords')下载原生停用词语料。

内容的提问来源于stack exchange,提问作者bdbrackett

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.27 19:09:18