NLTK停用词移除函数在循环中失效,求排查建议
问题排查与解决建议
一、核心逻辑错误:停用词判断顺序颠倒
你的removestopwords函数存在逻辑漏洞:先判断原词是否在全小写的停用词集合中,再转小写添加结果。比如原词是"The",停用词集合里只有"the",判断时会认为"The"不在集合内,最终把"the"加入结果,等于没移除停用词。你单独测试时可能用的是小写句子,所以看似正常,循环处理CSV里带大写的内容时就失效了。
修复后的函数代码:
def removestopwords(tokens): """This function removes the stop words from the tokens""" stop_words = set(stopwords.words("english")) filtered_tokens = list() for word in tokens: lower_word = word.lower() if lower_word not in stop_words: filtered_tokens.append(lower_word) return filtered_tokens
二、CSV行处理的潜在问题
- 跳过表头行:如果你的CSV第一行是表头(比如
label,headline),循环会把表头当成数据处理,导致结果异常。需添加跳过逻辑:
with open(SENTIMENT_CSV, 'rt', encoding = 'ISO-8859-1') as sobj: sdata = csv.reader(sobj) next(sdata) # 跳过表头行 all_tokenwords = list() tokenword_label = list() for row in sdata: # 后续代码保持不变
- 过滤无效行:CSV中可能存在空行或
row[1]为空的情况,导致分词无结果。添加判断跳过这类行:
for row in sdata: if len(row) < 2 or not row[1].strip(): continue # 跳过无内容或格式错误的行 tokens = tokenize(row[1]) filtered_tokens = removestopwords(tokens) all_tokenwords.extend(filtered_tokens) tokenword_label.append([filtered_tokens,row[0]])
三、分词函数优化(可选)
当前用split()分词会保留标点(比如"stock,"会被当成一个词),建议用NLTK专业分词工具更精准:
from nltk.tokenize import word_tokenize def tokenize(sentence): """ This function does the task of converting a sentence into a set of words""" t_words = word_tokenize(sentence) return t_words
使用前需下载分词模型:nltk.download('punkt')
内容的提问来源于stack exchange,提问作者axelalex15
相关产品推荐
相关产品推荐

