执行groupby遇TypeError: unhashable type: 'list',求分组统计均值并绘图
问题说明
我有一份科学问题相关的JSON数据集,单条数据结构如下:
[ { "Question": "What is the scientific method and why is it important?", "parameters": "a 7-11 year-old person", "Response": { "content": "It's a method for finding answers to questions", "analysis": { "text_num_chars": 63, "text_num_words": 8, "text_num_sents": 1 } } }, { "Question": "What is the difference between a theory and a hypothesis?", "parameters": "8th & 9th graders", "Response": { "content": "theory is an explanation of a phenomenon that is based on observations, experiments, and data.A hypothesis, on the other hand, is a testable statement about the natural world that can be used to build more complex hypotheses or theories.", "analysis": { "text_num_chars": 283, "text_num_words": 34, "text_num_sents": 2 } } } ]
将其转换为Pandas DataFrame后结构如下:
df = pd.DataFrame({'Questions': lst_questions, 'Parameters': lst_parameters, 'Metrics': lst_metrics})
其中df['Metrics']列的元素是字典,示例输出为:
[{'text_num_chars': 673, 'text_num_words': 100, 'text_num_sents': 5}, {'text_num_chars': 635, 'text_num_words': 101, 'text_num_sents': 7}]
尝试按Parameters列分组计算text_num_chars和text_num_words的均值时,触发错误:TypeError: unhashable type: 'list',最终需要绘制以参数值为X轴、对应text_num_chars均值为Y轴的图表。
解决步骤
1. 展开Metrics字典为独立列
由于Metrics列存储的是字典,无法直接参与分组计算,需先将其展开为单独的数值列:
import pandas as pd # 展开Metrics列的字典为DataFrame metrics_expanded = df['Metrics'].apply(pd.Series) # 合并到原DataFrame并删除原Metrics列 df = pd.concat([df.drop('Metrics', axis=1), metrics_expanded], axis=1)
2. 修复Parameters列的不可哈希类型
报错unhashable type: 'list'说明Parameters列中存在列表类型的数据,需将其转换为可哈希的字符串类型:
# 将列表类型的参数转为逗号分隔的字符串 df['Parameters'] = df['Parameters'].apply(lambda x: ', '.join(x) if isinstance(x, list) else x)
3. 分组计算均值
现在可以正常按Parameters分组,计算目标指标的均值:
# 按Parameters分组,计算text_num_chars和text_num_words的均值 grouped_stats = df.groupby('Parameters')[['text_num_chars', 'text_num_words']].mean().reset_index()
4. 绘制图表
使用Matplotlib绘制柱状图:
import matplotlib.pyplot as plt plt.figure(figsize=(10, 6)) # 绘制text_num_chars均值的柱状图 plt.bar(grouped_stats['Parameters'], grouped_stats['text_num_chars'], color='#1f77b4') # 设置图表标签与标题 plt.xlabel('目标受众') plt.ylabel('文本字符数均值') plt.title('不同受众对应的回答文本字符数均值') # 旋转X轴标签避免重叠 plt.xticks(rotation=45, ha='right') # 调整布局 plt.tight_layout() plt.show()
内容的提问来源于stack exchange,提问作者Donya
相关产品推荐
相关产品推荐

