Python列表转单词计数字典报unhashable type: list错误如何解决?
报错成因
- 核心问题出在
wordList = [cleanText.split()]这行代码:str.split()方法本身返回的就是存储所有单词的一维列表,外层额外加[]会把单词列表再嵌套一层,最终wordList的结构变为[[单词1, 单词2, 单词3...]] - 遍历
wordList时,拿到的item是内层的整个单词列表,而Python字典的键必须是可哈希的不可变类型,列表是可变类型,不允许作为字典键,因此触发unhashable type: list报错 - 额外问题:代码中导入了
collections.Counter但未使用,该工具可以直接实现单词计数需求,不需要手动写循环统计。
修复方案
有两种可选修改方式,都可以实现单词计数需求:
方式1:保留手动计数逻辑,修复列表嵌套问题
仅需删除cleanText.split()外层的多余方括号即可,修改后完整代码为:
import re file = open("dialog.txt") fileContents = file.read().lower() file.close() cleanupText = re.compile("[^a-zA-Z]") cleanText = re.sub(cleanupText," ",fileContents) # 去掉外层[],直接取split返回的单词列表 wordList = cleanText.split() wordDictionary = {} for item in wordList: # 不需要调用.keys(),直接判断元素是否在字典的键中即可 if item in wordDictionary: wordDictionary[item] += 1 else: wordDictionary[item] = 1 print(wordDictionary)
方式2:使用已导入的Counter工具简化代码
Counter原生支持可迭代对象的元素计数,无需手动写循环,代码更简洁:
import re from collections import Counter file = open("dialog.txt") fileContents = file.read().lower() file.close() cleanupText = re.compile("[^a-zA-Z]") cleanText = re.sub(cleanupText," ",fileContents) wordList = cleanText.split() # 直接用Counter生成计数字典,Counter是dict的子类,可直接当做字典使用 wordDictionary = Counter(wordList) print(dict(wordDictionary)) # 转成普通字典输出也可以,直接打印Counter也能看到计数结果
内容的提问来源于stack exchange,提问作者Jacob Thacker
相关产品推荐
相关产品推荐

