You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

NLTK停用词移除函数在循环中失效,求排查建议

问题排查与解决建议

一、核心逻辑错误:停用词判断顺序颠倒

你的removestopwords函数存在逻辑漏洞:先判断原词是否在全小写的停用词集合中,再转小写添加结果。比如原词是"The",停用词集合里只有"the",判断时会认为"The"不在集合内,最终把"the"加入结果,等于没移除停用词。你单独测试时可能用的是小写句子,所以看似正常,循环处理CSV里带大写的内容时就失效了。

修复后的函数代码:

def removestopwords(tokens):
    """This function removes the stop words from the tokens"""
    stop_words = set(stopwords.words("english"))
    filtered_tokens = list()
    for word in tokens:
        lower_word = word.lower()
        if lower_word not in stop_words:
            filtered_tokens.append(lower_word)
    return filtered_tokens

二、CSV行处理的潜在问题

  1. 跳过表头行:如果你的CSV第一行是表头(比如label,headline),循环会把表头当成数据处理,导致结果异常。需添加跳过逻辑:
with open(SENTIMENT_CSV, 'rt', encoding = 'ISO-8859-1') as sobj:
    sdata = csv.reader(sobj)
    next(sdata)  # 跳过表头行
    all_tokenwords = list()
    tokenword_label = list()
    for row in sdata: 
        # 后续代码保持不变
  1. 过滤无效行:CSV中可能存在空行或row[1]为空的情况,导致分词无结果。添加判断跳过这类行:
for row in sdata: 
    if len(row) < 2 or not row[1].strip():
        continue  # 跳过无内容或格式错误的行
    tokens = tokenize(row[1])
    filtered_tokens = removestopwords(tokens)
    all_tokenwords.extend(filtered_tokens)
    tokenword_label.append([filtered_tokens,row[0]])

三、分词函数优化(可选)

当前用split()分词会保留标点(比如"stock,"会被当成一个词),建议用NLTK专业分词工具更精准:

from nltk.tokenize import word_tokenize

def tokenize(sentence):
    """ This function does the task of converting a sentence into a set of words"""
    t_words = word_tokenize(sentence)
    return t_words

使用前需下载分词模型:nltk.download('punkt')

内容的提问来源于stack exchange,提问作者axelalex15

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 09:15:32