如何在pandas DataFrame列处理中应用自定义扩展停用词表
修改方法
你已经完成了自定义停用词和nltk原生停用词的集合合并,只需要把停用词过滤逻辑里引用的停用词变量,从原来的stop_words替换为合并后的new_stopwords_list即可,原有分词过滤的逻辑完全不需要调整。
完整可运行代码
from nltk.corpus import stopwords import pandas as pd # 加载nltk原生英文停用词表 stop_words = set(stopwords.words('english')) # 自定义补充停用词列表 new_stopwords = ['satisfying', 'satisfy', 'satisfied', 'clemson', 'university', 'institution', 'disappointing', 'disappoint', 'disappointed', 'experience', 'would', 'should'] # 合并生成扩展停用词集合 new_stopwords_list = stop_words.union(new_stopwords) # 过滤停用词,生成新列 df['stopwords_removed'] = df['no_punc'].apply(lambda x: [word for word in x if word not in new_stopwords_list]) df.head()
注意事项
- 你的前置预处理已经把
no_punc列的文本完成了分词、全小写转换,当前自定义停用词列表内的词均为小写格式,和待过滤文本格式完全匹配,不会出现漏匹配的问题。 - 如果后续需要新增停用词,直接将对应小写格式的词汇加入
new_stopwords列表即可,不需要修改后续过滤逻辑。 - 如果首次运行nltk相关代码报错,可提前执行
nltk.download('stopwords')下载原生停用词语料。
内容的提问来源于stack exchange,提问作者bdbrackett
相关产品推荐
相关产品推荐

