如何从DataFrame中移除文件内的停用词?代码未生效求助
移除DataFrame句子停用词无效的问题解决
问题场景
我有一个存储停用词的txt文件,想要移除tokenized_tweets列表里的停用词,但运行以下代码后没有效果,停用词依然存在:
f = open("stopwords.txt", "r") stopword_list = [] for line in f: stripped_line = line.strip() line_list = stripped_line.split() stopword_list.append(line_list[0]) f.close() len(stopword_list) tokens_without_sw = [word for word in tokenized_tweets if not word in stopword_list] print("After stopwords removed") print(tokens_without_sw)
可能原因及解决方法
1. 停用词加载不规范
原代码用split()后取line_list[0]属于多余操作(如果停用词文件每行仅一个词),且未处理空行、大小写差异问题。优化加载逻辑:
# 加载停用词,统一转小写并跳过空行 stopword_list = [] with open("stopwords.txt", "r", encoding="utf-8") as f: for line in f: word = line.strip() if word: # 跳过空行 stopword_list.append(word.lower()) # 转成集合,大幅提升查找效率 stopword_set = set(stopword_list)
2. 词形不匹配
tokenized_tweets中的词和停用词可能存在大小写差异(比如停用词是the,token是The),或token附带标点(比如hello,),导致匹配失败。处理方式:
import string # 先清洗token:去除标点、统一转小写 clean_tokens = [] for word in tokenized_tweets: cleaned_word = word.strip(string.punctuation).lower() if cleaned_word: # 过滤空字符串 clean_tokens.append(cleaned_word) # 过滤停用词 tokens_without_sw = [word for word in clean_tokens if word not in stopword_set]
3. 文件路径错误
如果stopwords.txt不在当前脚本的工作目录下,会加载出空的停用词列表。可以先打印len(stopword_list)确认长度,若为0则改用绝对路径打开文件:
with open("C:/your/actual/path/stopwords.txt", "r", encoding="utf-8") as f: # 后续加载逻辑不变
4. 核对内容交集
先打印tokenized_tweets和stopword_list的具体内容,确认两者是否真的存在交集。比如如果token是分词后的完整词汇,而停用词是单个字,自然无法匹配。
内容的提问来源于stack exchange,提问作者Zulfi A
相关产品推荐
相关产品推荐

