You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python生成词云遇TypeError: unhashable type: 'list'及代码排查

问题

尝试基于pandas DataFrame生成词云,通过for循环去除标点符号与停用词时遇到TypeError: unhashable type: 'list'错误,不清楚错误含义。同时希望排查所有潜在问题,且要求用for循环实现(已知列表推导式可简化,但暂不使用)。

数据示例

description
This tremendous 100% varietal wine hails from ...
Ripe aromas of fig, blackberry and cassis are ...
This spent 20 months in 30% new French oak, an...

出错代码

punctuations = '''!()-[]{};:'"\,<>./?@#$%^&*_~'''
uninteresting_words = ["the", "a", "to", "if", "is", "it", "of", "and", "or", "an", "as", "i", "me", "my", \
    "we", "our", "ours", "you", "your", "yours", "he", "she", "him", "his", "her", "hers", "its", "they", "them", \
    "their", "what", "which", "who", "whom", "this", "that", "am", "are", "was", "were", "be", "been", "being", \
    "have", "has", "had", "do", "does", "did", "but", "at", "by", "with", "from", "here", "when", "where", "how", \
    "all", "any", "both", "each", "few", "more", "some", "such", "no", "nor", "too", "very", "can", "will", "j"]


def wordcloudfunc(data):
  frequency_count = {}
  refined_text = ""

  for p in punctuations:
    data = data.replace(p,"")

  refined_text = data.str.split()

  for word in refined_text:
    if word not in uninteresting_words:
      frequency_count[word] += 1 # this is where the error occur
    else:
      frequency_count[word] = 1

  #wordcloud
    cloud = wordcloud.WordCloud()
    cloud.generate_from_frequencies(frequency_count)
    return cloud.to_array()

myimage = wordcloudfunc(df['description'])
plt.imshow(myimage, interpolation = 'nearest')
plt.axis('off')
plt.show()

问题分析与修复

1. TypeError: unhashable type: 'list' 错误原因

data.str.split() 返回的是Series对象,每个元素是一个单词列表(比如第一行文本会被拆成['This', 'tremendous', ...])。你直接遍历refined_text时,word变量拿到的是整个列表,而字典的键必须是可哈希类型(比如字符串),列表是不可哈希的,因此触发报错。

2. 其他潜在问题

  • 标点替换逻辑错误:data.replace(p,"") 对Series使用时默认是全匹配替换,应该用data.str.replace(p, "")才能按单个字符替换标点。
  • 大小写不一致:原文本中的单词(如"This")未转小写,会和停用词里的"this"被当成不同单词,导致统计重复。
  • 频率统计逻辑完全颠倒:当前代码中,非停用词执行累加、停用词执行初始化,实际应该只统计非停用词。同时首次出现的单词直接累加会触发KeyError。
  • 词云代码缩进错误:cloud = wordcloud.WordCloud()缩进在for循环内,会每次循环创建新对象且提前return,导致只处理第一个文本行就结束。
  • 缺失必要导入:代码用到wordcloud和plt但未导入相关模块。

3. 修复后的代码

import pandas as pd
from wordcloud import WordCloud
import matplotlib.pyplot as plt

punctuations = '''!()-[]{};:'"\,<>./?@#$%^&*_~'''
uninteresting_words = ["the", "a", "to", "if", "is", "it", "of", "and", "or", "an", "as", "i", "me", "my", \
    "we", "our", "ours", "you", "your", "yours", "he", "she", "him", "his", "her", "hers", "its", "they", "them", \
    "their", "what", "which", "who", "whom", "this", "that", "am", "are", "was", "were", "be", "been", "being", \
    "have", "has", "had", "do", "does", "did", "but", "at", "by", "with", "from", "here", "when", "where", "how", \
    "all", "any", "both", "each", "few", "more", "some", "such", "no", "nor", "too", "very", "can", "will", "j"]


def wordcloudfunc(data):
    frequency_count = {}
    
    # 1. 遍历标点,逐个替换去除
    for p in punctuations:
        data = data.str.replace(p, "")
    
    # 2. 遍历每一行的单词列表
    for words_list in data.str.split():
        # 遍历列表中的每个单词
        for word in words_list:
            # 转小写统一格式
            lower_word = word.lower()
            # 仅统计非停用词
            if lower_word not in uninteresting_words:
                # 处理首次出现的单词,避免KeyError
                if lower_word in frequency_count:
                    frequency_count[lower_word] += 1
                else:
                    frequency_count[lower_word] = 1
    
    # 3. 生成词云并返回
    cloud = WordCloud()
    cloud.generate_from_frequencies(frequency_count)
    return cloud.to_array()

# 假设df为你的目标DataFrame
myimage = wordcloudfunc(df['description'])
plt.imshow(myimage, interpolation='nearest')
plt.axis('off')
plt.show()

修复说明

  • 调整遍历逻辑:先遍历Series中的每个单词列表,再遍历列表内的单个单词,确保word为字符串类型,解决哈希错误。
  • 替换标点方法:用data.str.replace实现逐字符替换,正确去除标点。
  • 统一大小写:将所有单词转小写,避免因大小写导致的重复统计。
  • 修正频率统计逻辑:仅统计非停用词,且处理首次出现的单词,避免KeyError。
  • 调整缩进:将词云生成代码移到循环外,确保统计完所有单词后再生成词云。
  • 添加必要的模块导入语句,保证代码可运行。

内容的提问来源于stack exchange,提问作者Data Beginner

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.17 13:50:24