You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

执行groupby遇TypeError: unhashable type: 'list',求分组统计均值并绘图

问题说明

我有一份科学问题相关的JSON数据集,单条数据结构如下:

[
    {
        "Question": "What is the scientific method  and why is it important?",
        "parameters": "a 7-11 year-old person",
        "Response": {
            "content": "It's a method for finding answers to questions",
            "analysis": {
                "text_num_chars": 63,
                "text_num_words": 8,
                "text_num_sents": 1
            }
        }
    },
    {
        "Question": "What is the difference between a theory and a hypothesis?",
        "parameters": "8th & 9th graders",
        "Response": {
            "content": "theory is an explanation of a phenomenon that is based on observations, experiments, and data.A hypothesis, on the other hand, is a testable statement about the natural world that can be used to build more complex hypotheses or theories.",
            "analysis": {
                "text_num_chars": 283,
                "text_num_words": 34,
                "text_num_sents": 2
            }
        }
    }
]

将其转换为Pandas DataFrame后结构如下:

df = pd.DataFrame({'Questions': lst_questions, 
                   'Parameters': lst_parameters,
                   'Metrics': lst_metrics})

其中df['Metrics']列的元素是字典,示例输出为:

[{'text_num_chars': 673, 'text_num_words': 100, 'text_num_sents': 5}, {'text_num_chars': 635, 'text_num_words': 101, 'text_num_sents': 7}]

尝试按Parameters列分组计算text_num_chars和text_num_words的均值时,触发错误:TypeError: unhashable type: 'list',最终需要绘制以参数值为X轴、对应text_num_chars均值为Y轴的图表。

解决步骤

1. 展开Metrics字典为独立列

由于Metrics列存储的是字典,无法直接参与分组计算,需先将其展开为单独的数值列:

import pandas as pd

# 展开Metrics列的字典为DataFrame
metrics_expanded = df['Metrics'].apply(pd.Series)
# 合并到原DataFrame并删除原Metrics列
df = pd.concat([df.drop('Metrics', axis=1), metrics_expanded], axis=1)

2. 修复Parameters列的不可哈希类型

报错unhashable type: 'list'说明Parameters列中存在列表类型的数据,需将其转换为可哈希的字符串类型:

# 将列表类型的参数转为逗号分隔的字符串
df['Parameters'] = df['Parameters'].apply(lambda x: ', '.join(x) if isinstance(x, list) else x)

3. 分组计算均值

现在可以正常按Parameters分组,计算目标指标的均值:

# 按Parameters分组,计算text_num_chars和text_num_words的均值
grouped_stats = df.groupby('Parameters')[['text_num_chars', 'text_num_words']].mean().reset_index()

4. 绘制图表

使用Matplotlib绘制柱状图:

import matplotlib.pyplot as plt

plt.figure(figsize=(10, 6))
# 绘制text_num_chars均值的柱状图
plt.bar(grouped_stats['Parameters'], grouped_stats['text_num_chars'], color='#1f77b4')
# 设置图表标签与标题
plt.xlabel('目标受众')
plt.ylabel('文本字符数均值')
plt.title('不同受众对应的回答文本字符数均值')
# 旋转X轴标签避免重叠
plt.xticks(rotation=45, ha='right')
# 调整布局
plt.tight_layout()
plt.show()

内容的提问来源于stack exchange,提问作者Donya

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.28 19:35:01