如何在Pandas DataFrame中统计列值顺序无关的重复行?
可行!按值集合分组统计重复行的实现方案
完全可以实现这个需求,核心思路是先把每行的列值转换成排序后的统一标识,让值集合相同但顺序不同的行拥有相同的分组键,再基于这个键统计出现次数。以下是具体实现步骤:
1. 构造示例数据
import pandas as pd df = pd.DataFrame({ 'fruit1': ['apple', 'cherry', 'apple', 'banana'], 'fruit2': ['banana', 'orange', 'banana', 'apple'] })
2. 生成统一分组键
对每行的列值进行排序,转成元组(元组可哈希,能作为分组依据):
df['group_key'] = df.apply(lambda row: tuple(sorted(row)), axis=1)
3. 分组统计次数
按分组键聚合,统计每组的出现次数:
counts = df.groupby('group_key').size().reset_index(name='occurences')
4. 还原列格式并整理结果
将分组键拆回原列,删除临时列后调整顺序:
counts[['fruit1', 'fruit2']] = pd.DataFrame(counts['group_key'].tolist(), index=counts.index) result = counts.drop('group_key', axis=1)[['fruit1', 'fruit2', 'occurences']]
最终输出结果和你期望的一致:
fruit1 fruit2 occurences 0 apple banana 3 1 cherry orange 1
简化版实现(无需临时列)
可以把步骤合并,用value_counts直接统计:
result = ( df.apply(lambda row: tuple(sorted(row)), axis=1) .value_counts() .reset_index(name='occurences') ) result[['fruit1', 'fruit2']] = pd.DataFrame(result['index'].tolist(), index=result.index) result = result.drop('index', axis=1)[['fruit1', 'fruit2', 'occurences']]
原理说明
通过排序让apple+banana和banana+apple生成相同的分组键('apple', 'banana'),这样就能把它们归为同一组统计次数,完美解决了列值顺序不同但值集合相同的重复行统计问题。
内容的提问来源于stack exchange,提问作者Clabis
相关产品推荐
相关产品推荐

