Python中Pandas文本列停用词未被成功删除的问题排查
问题分析与解决方案
嘿,我一眼就看出你代码里的逻辑问题啦!咱们来拆解一下:
你的代码为什么没生效?
你写的这段列表推导式:
train['text'] = [i for i in train['text'].str.split() if i not in stopword]
这里的i是整个文本的分词列表(比如["今天", "是", "周一"]),而你的stopword是单个停用词的列表(比如["是"])。你现在在判断「整个分词列表是否属于停用词列表」,这显然永远是True——因为停用词里根本没有列表元素,所以所有内容都被原封不动保留了,停用词自然还在文本里。
正确的解决方法
咱们需要对每个分词列表里的单个词进行过滤,而不是拿整个列表去匹配。另外,把停用词转成set能大幅提升查询效率(特别是停用词数量多的时候),这是个实用小技巧哦。
步骤1:优化停用词存储
# 加载停用词并转成集合 stopword = [line.rstrip() for line in open('gdrive/My Drive/stopword.txt')] stopword_set = set(stopword) # 集合的成员查询比列表快N倍
步骤2:过滤每个文本中的停用词
根据你的需求,有两种处理方式:
- 如果希望保留分词后的列表格式:
train['text'] = train['text'].str.split().apply(lambda tokens: [word for word in tokens if word not in stopword_set])
- 如果希望把过滤后的词重新拼成字符串(恢复成文本格式):
train['text'] = train['text'].str.split().apply(lambda tokens: ' '.join([word for word in tokens if word not in stopword_set]))
额外提醒
如果你的文本存在大小写不一致的情况(比如停用词是小写"is",但文本里是大写"Is"),记得先统一大小写:
train['text'] = train['text'].str.lower().str.split().apply(lambda tokens: ' '.join([word for word in tokens if word not in stopword_set]))
内容的提问来源于stack exchange,提问作者Newbie Student
相关产品推荐
相关产品推荐

