You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于列表内容(忽略顺序)分组DataFrame并按组大小排序的报错问题

解决方法:按列表元素(忽略顺序)分组并统计

错误原因

collections.Counter是可变容器类型,不具备哈希特性,无法直接作为groupby的分组键,因此会抛出TypeError: unhashable type: 'Counter'。

可行方案

方案1:排序后转元组(通用且高效)

将列表元素排序后转为元组(元组是不可变可哈希类型),这样无论原列表元素顺序如何,相同元素组合都会生成一致的分组键:

import pandas as pd

d1 = {'id': ["car", "car", "bus", "plane", "plane"], 'value': [["a","b"], ["b","a"], ["a","b"], ["c","d"], ["d","c"]]}
df1 = pd.DataFrame(data=d1)

# 生成排序后的元组作为分组键
df1['group_key'] = df1['value'].apply(lambda x: tuple(sorted(x)))
# 分组统计并按组大小升序排序
result = df1.groupby('group_key').size().sort_values(ascending=True)
print(result)

输出结果:

group_key
(a, b)    3
(c, d)    2
dtype: int64

方案2:将Counter转为可哈希结构(适合含重复元素的列表)

如果列表存在重复元素(如["a","a","b"]),可以将Counter的键值对排序后转为元组,确保相同元素计数的列表生成一致的分组键:

import pandas as pd
from collections import Counter

d1 = {'id': ["car", "car", "bus", "plane", "plane"], 'value': [["a","b"], ["b","a"], ["a","b"], ["c","d"], ["d","c"]]}
df1 = pd.DataFrame(data=d1)

# 将Counter转换为排序后的元组,使其可哈希
df1['group_key'] = df1['value'].apply(lambda x: tuple(sorted(Counter(x).items())))
result = df1.groupby('group_key').size().sort_values(ascending=True)
print(result)

输出结果与方案1一致,同时支持含重复元素的列表分组。

内容的提问来源于stack exchange,提问作者Limmi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.26 11:05:22