You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从DataFrame中移除文件内的停用词?代码未生效求助

移除DataFrame句子停用词无效的问题解决

问题场景

我有一个存储停用词的txt文件,想要移除tokenized_tweets列表里的停用词,但运行以下代码后没有效果,停用词依然存在:

f = open("stopwords.txt", "r")
stopword_list = []
for line in f:
    stripped_line = line.strip()
    line_list = stripped_line.split()
    stopword_list.append(line_list[0])
f.close()

len(stopword_list)

tokens_without_sw = [word for word in tokenized_tweets if not word in stopword_list]
print("After stopwords removed")
print(tokens_without_sw)

可能原因及解决方法

1. 停用词加载不规范

原代码用split()后取line_list[0]属于多余操作(如果停用词文件每行仅一个词),且未处理空行、大小写差异问题。优化加载逻辑:

# 加载停用词,统一转小写并跳过空行
stopword_list = []
with open("stopwords.txt", "r", encoding="utf-8") as f:
    for line in f:
        word = line.strip()
        if word:  # 跳过空行
            stopword_list.append(word.lower())
# 转成集合,大幅提升查找效率
stopword_set = set(stopword_list)

2. 词形不匹配

tokenized_tweets中的词和停用词可能存在大小写差异(比如停用词是the,token是The),或token附带标点(比如hello,),导致匹配失败。处理方式:

import string

# 先清洗token:去除标点、统一转小写
clean_tokens = []
for word in tokenized_tweets:
    cleaned_word = word.strip(string.punctuation).lower()
    if cleaned_word:  # 过滤空字符串
        clean_tokens.append(cleaned_word)

# 过滤停用词
tokens_without_sw = [word for word in clean_tokens if word not in stopword_set]

3. 文件路径错误

如果stopwords.txt不在当前脚本的工作目录下,会加载出空的停用词列表。可以先打印len(stopword_list)确认长度,若为0则改用绝对路径打开文件:

with open("C:/your/actual/path/stopwords.txt", "r", encoding="utf-8") as f:
    # 后续加载逻辑不变

4. 核对内容交集

先打印tokenized_tweets和stopword_list的具体内容,确认两者是否真的存在交集。比如如果token是分词后的完整词汇,而停用词是单个字,自然无法匹配。

内容的提问来源于stack exchange,提问作者Zulfi A

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.31 03:27:39