Python生成词云遇TypeError: unhashable type: 'list'及代码排查
问题
尝试基于pandas DataFrame生成词云,通过for循环去除标点符号与停用词时遇到TypeError: unhashable type: 'list'错误,不清楚错误含义。同时希望排查所有潜在问题,且要求用for循环实现(已知列表推导式可简化,但暂不使用)。
数据示例
| description |
|---|
| This tremendous 100% varietal wine hails from ... |
| Ripe aromas of fig, blackberry and cassis are ... |
| This spent 20 months in 30% new French oak, an... |
出错代码
punctuations = '''!()-[]{};:'"\,<>./?@#$%^&*_~''' uninteresting_words = ["the", "a", "to", "if", "is", "it", "of", "and", "or", "an", "as", "i", "me", "my", \ "we", "our", "ours", "you", "your", "yours", "he", "she", "him", "his", "her", "hers", "its", "they", "them", \ "their", "what", "which", "who", "whom", "this", "that", "am", "are", "was", "were", "be", "been", "being", \ "have", "has", "had", "do", "does", "did", "but", "at", "by", "with", "from", "here", "when", "where", "how", \ "all", "any", "both", "each", "few", "more", "some", "such", "no", "nor", "too", "very", "can", "will", "j"] def wordcloudfunc(data): frequency_count = {} refined_text = "" for p in punctuations: data = data.replace(p,"") refined_text = data.str.split() for word in refined_text: if word not in uninteresting_words: frequency_count[word] += 1 # this is where the error occur else: frequency_count[word] = 1 #wordcloud cloud = wordcloud.WordCloud() cloud.generate_from_frequencies(frequency_count) return cloud.to_array() myimage = wordcloudfunc(df['description']) plt.imshow(myimage, interpolation = 'nearest') plt.axis('off') plt.show()
问题分析与修复
1. TypeError: unhashable type: 'list' 错误原因
data.str.split() 返回的是Series对象,每个元素是一个单词列表(比如第一行文本会被拆成['This', 'tremendous', ...])。你直接遍历refined_text时,word变量拿到的是整个列表,而字典的键必须是可哈希类型(比如字符串),列表是不可哈希的,因此触发报错。
2. 其他潜在问题
- 标点替换逻辑错误:
data.replace(p,"")对Series使用时默认是全匹配替换,应该用data.str.replace(p, "")才能按单个字符替换标点。 - 大小写不一致:原文本中的单词(如"This")未转小写,会和停用词里的"this"被当成不同单词,导致统计重复。
- 频率统计逻辑完全颠倒:当前代码中,非停用词执行累加、停用词执行初始化,实际应该只统计非停用词。同时首次出现的单词直接累加会触发
KeyError。 - 词云代码缩进错误:
cloud = wordcloud.WordCloud()缩进在for循环内,会每次循环创建新对象且提前return,导致只处理第一个文本行就结束。 - 缺失必要导入:代码用到
wordcloud和plt但未导入相关模块。
3. 修复后的代码
import pandas as pd from wordcloud import WordCloud import matplotlib.pyplot as plt punctuations = '''!()-[]{};:'"\,<>./?@#$%^&*_~''' uninteresting_words = ["the", "a", "to", "if", "is", "it", "of", "and", "or", "an", "as", "i", "me", "my", \ "we", "our", "ours", "you", "your", "yours", "he", "she", "him", "his", "her", "hers", "its", "they", "them", \ "their", "what", "which", "who", "whom", "this", "that", "am", "are", "was", "were", "be", "been", "being", \ "have", "has", "had", "do", "does", "did", "but", "at", "by", "with", "from", "here", "when", "where", "how", \ "all", "any", "both", "each", "few", "more", "some", "such", "no", "nor", "too", "very", "can", "will", "j"] def wordcloudfunc(data): frequency_count = {} # 1. 遍历标点,逐个替换去除 for p in punctuations: data = data.str.replace(p, "") # 2. 遍历每一行的单词列表 for words_list in data.str.split(): # 遍历列表中的每个单词 for word in words_list: # 转小写统一格式 lower_word = word.lower() # 仅统计非停用词 if lower_word not in uninteresting_words: # 处理首次出现的单词,避免KeyError if lower_word in frequency_count: frequency_count[lower_word] += 1 else: frequency_count[lower_word] = 1 # 3. 生成词云并返回 cloud = WordCloud() cloud.generate_from_frequencies(frequency_count) return cloud.to_array() # 假设df为你的目标DataFrame myimage = wordcloudfunc(df['description']) plt.imshow(myimage, interpolation='nearest') plt.axis('off') plt.show()
修复说明
- 调整遍历逻辑:先遍历Series中的每个单词列表,再遍历列表内的单个单词,确保
word为字符串类型,解决哈希错误。 - 替换标点方法:用
data.str.replace实现逐字符替换,正确去除标点。 - 统一大小写:将所有单词转小写,避免因大小写导致的重复统计。
- 修正频率统计逻辑:仅统计非停用词,且处理首次出现的单词,避免
KeyError。 - 调整缩进:将词云生成代码移到循环外,确保统计完所有单词后再生成词云。
- 添加必要的模块导入语句,保证代码可运行。
内容的提问来源于stack exchange,提问作者Data Beginner
相关产品推荐
相关产品推荐

