嵌套字典转Pandas DataFrame调用describe()报TypeError: unhashable type: 'list'
解决嵌套字典转Pandas DataFrame及
describe()报错问题 嘿,我来帮你搞定这个问题!先给你拆解清楚错误原因,再给你对应解决方案。
为什么会报错?
你遇到的TypeError: unhashable type: 'list',根源在于你用pd.DataFrame.from_dict(dictionary, orient='index')创建的DataFrame里,a、b、c这些列的单元格值都是列表类型。而describe()方法在处理这类列时,会尝试调用value_counts()统计元素出现次数,但列表是不可哈希的对象——哈希表没法处理它,所以直接抛出了错误。
举个直观的例子,你创建的DataFrame实际结构是这样的:
| a | b | c | |
|---|---|---|---|
| A | ['1','2','3'] | ['4','5'] | ['6'] |
| B | ['7'] | ['8','9'] | NaN |
这种存列表的列,自然没法用describe()做常规统计。
解决方案(分三种场景)
根据你想要的最终DataFrame格式,我给你三种不同的处理方式:
场景1:把列表转成字符串,让describe()正常工作
如果你只是想让describe()能运行,同时保留每个列的多值信息,可以把列表转换成字符串(可哈希类型):
import pandas as pd # 你的原始字典 dictionary = {'A' : {'a' : ['1', '2', '3'], 'b' : ['4', '5'], 'c' : ['6']}, 'B' : {'a' : ['7'], 'b' : ['8', '9']}} # 遍历字典,把所有列表转成逗号分隔的字符串 processed_dict = { outer_key: {inner_key: ', '.join(inner_values) for inner_key, inner_values in inner_dict.items()} for outer_key, inner_dict in dictionary.items() } # 创建DataFrame并填充空值为空白字符串 df = pd.DataFrame.from_dict(processed_dict, orient='index').fillna('') print(df) # 现在可以正常调用describe()了 print(df.describe())
生成的DataFrame是这样的:
| a | b | c | |
|---|---|---|---|
| A | 1, 2, 3 | 4, 5 | 6 |
| B | 7 | 8, 9 |
场景2:把列表展开成多行(更适合数据分析)
如果你希望数据结构更规范,方便后续分析,建议把列表里的每个元素拆成单独的行:
import pandas as pd dictionary = {'A' : {'a' : ['1', '2', '3'], 'b' : ['4', '5'], 'c' : ['6']}, 'B' : {'a' : ['7'], 'b' : ['8', '9']}} # 先把嵌套字典扁平化 flat_data = [] for outer_key, inner_dict in dictionary.items(): for inner_key, values in inner_dict.items(): for val in values: flat_data.append({ 'group': outer_key, 'category': inner_key, 'value': val }) # 转成宽格式DataFrame df = pd.DataFrame(flat_data).pivot(index='group', columns='category', values='value') print(df)
输出的结构化DataFrame:
| group | a | b | c |
|---|---|---|---|
| A | 1 | 4 | 6 |
| A | 2 | 5 | NaN |
| A | 3 | NaN | NaN |
| B | 7 | 8 | NaN |
| B | NaN | 9 | NaN |
这种格式下describe()可以直接正常运行,完全满足数据分析的需求。
场景3:完全匹配你要的格式(单元格存列表)
如果你就想保留单元格里的列表(和你给出的期望格式一致),那直接用你原来的代码就行,但要注意describe()会报错——如果需要统计,就得按场景1的方法转成字符串:
import pandas as pd dictionary = {'A' : {'a' : ['1', '2', '3'], 'b' : ['4', '5'], 'c' : ['6']}, 'B' : {'a' : ['7'], 'b' : ['8', '9']}} df = pd.DataFrame.from_dict(dictionary, orient='index').fillna('') print(df)
输出和你期望的一致:
| a | b | c | |
|---|---|---|---|
| A | ['1','2','3'] | ['4','5'] | ['6'] |
| B | ['7'] | ['8','9'] |
但此时调用df.describe()还是会报错,所以如果需要统计功能,一定要先把列表转成可哈希的类型哦。
内容的提问来源于stack exchange,提问作者SangHyeon Na
相关产品推荐
相关产品推荐

