基于列表内容(忽略顺序)分组DataFrame并按组大小排序的报错问题
解决方法:按列表元素(忽略顺序)分组并统计
错误原因
collections.Counter是可变容器类型,不具备哈希特性,无法直接作为groupby的分组键,因此会抛出TypeError: unhashable type: 'Counter'。
可行方案
方案1:排序后转元组(通用且高效)
将列表元素排序后转为元组(元组是不可变可哈希类型),这样无论原列表元素顺序如何,相同元素组合都会生成一致的分组键:
import pandas as pd d1 = {'id': ["car", "car", "bus", "plane", "plane"], 'value': [["a","b"], ["b","a"], ["a","b"], ["c","d"], ["d","c"]]} df1 = pd.DataFrame(data=d1) # 生成排序后的元组作为分组键 df1['group_key'] = df1['value'].apply(lambda x: tuple(sorted(x))) # 分组统计并按组大小升序排序 result = df1.groupby('group_key').size().sort_values(ascending=True) print(result)
输出结果:
group_key (a, b) 3 (c, d) 2 dtype: int64
方案2:将Counter转为可哈希结构(适合含重复元素的列表)
如果列表存在重复元素(如["a","a","b"]),可以将Counter的键值对排序后转为元组,确保相同元素计数的列表生成一致的分组键:
import pandas as pd from collections import Counter d1 = {'id': ["car", "car", "bus", "plane", "plane"], 'value': [["a","b"], ["b","a"], ["a","b"], ["c","d"], ["d","c"]]} df1 = pd.DataFrame(data=d1) # 将Counter转换为排序后的元组,使其可哈希 df1['group_key'] = df1['value'].apply(lambda x: tuple(sorted(Counter(x).items()))) result = df1.groupby('group_key').size().sort_values(ascending=True) print(result)
输出结果与方案1一致,同时支持含重复元素的列表分组。
内容的提问来源于stack exchange,提问作者Limmi
相关产品推荐
相关产品推荐

