Python代码出现NameError:'stop_words'未定义问题排查
问题分析与修复
错误根源
你在preprocess_text函数内部定义的stop_words是局部变量,仅在函数执行范围内有效,函数外部的代码(比如生成词云的部分)无法访问这个变量,因此触发NameError。
修复步骤
1. 将停用词定义改为全局变量
把stop_words的定义移到函数外部,让整个代码都能调用:
# 全局定义停用词集合 stop_words = set(stopwords.words("english")) stop_words.update(["a", "an", "and", "are", "as", "at", "be", "by", "for", "from", "has", "he", "in", "is", "it", "its", "of", "on", "that", "the", "to", "was", "were", "will", "with"]) # 预处理函数(直接使用全局的stop_words) def preprocess_text(text): text = text.lower() text = re.sub(r'\[^\\w\\s\]', '', text) words = word_tokenize(text) words = [word for word in words if word not in stop_words] lemmatizer = WordNetLemmatizer() words = [lemmatizer.lemmatize(word) for word in words] return words
2. 修正缩进错误
你统计正负特征的循环里存在缩进问题,会导致语法错误,需要调整:
positive_features = {} negative_features = {} for i in range(len(comments)): comment = comments[i] sentiment = sentiments[i] words = preprocess_text(comment) for word in words: # 此处需缩进,与外层for对齐 if sentiment > 0: if word in positive_features: positive_features[word] += 1 else: positive_features[word] = 1 elif sentiment < 0: if word in negative_features: negative_features[word] += 1 else: negative_features[word] = 1
3. 词云部分正常使用全局stop_words
修改后生成词云的代码就能正常调用stop_words,无需额外改动。
额外优化说明
将停用词设为全局变量,不仅解决了变量访问问题,还避免了每次调用预处理函数时重复创建相同的集合,提升了代码效率,同时保证函数内外使用的停用词完全一致,逻辑更统一。
内容的提问来源于stack exchange,提问作者cody
相关产品推荐
相关产品推荐

