Python Pandas DataFrame跨父分组查找所有子列重复值的实现方法
Pandas跨父分组查找子列重复值解决方案
实现逻辑
- 先将多层分组索引转为普通列,方便按列统计
- 自动识别子列,适配任意子列数量场景
- 逐个统计每个子列中值对应的不同父分组数量,筛选出跨父分组出现的重复值
- 合并所有符合条件的行,恢复原多层索引格式输出
完整代码
import pandas as pd # 构造示例数据 data = {'buildings': {0: 'mansion', 1: 'mansion', 2: 'house', 3: 'house', 4: 'house', 5: 'apartment', 6: 'apartment', 7: 'apartment', 8: 'apartment', 9: 'apartment', 10: 'apartment', 11: 'apartment', 12: 'condo', 13: 'condo', 14: 'condo', 15: 'condo', 16: 'condo', 17: 'condo'}, 'vehicles': {0: 'plane', 1: 'boat', 2: 'small car', 3: 'small car', 4: 'big car', 5: 'small truck', 6: 'big truck', 7: 'big truck', 8: 'big truck', 9: 'big truck', 10: 'big truck', 11: 'big truck', 12: 'condo', 13: 'condo', 14: 'condo', 15: 'condo', 16: 'condo', 17: 'condo'}, 'animals': {0: 'plane', 1: 'boat', 2: 'ape', 3: 'fish', 4: 'big car', 5: 'small truck', 6: 'chimp', 7: 'monkey', 8: 'lemur', 9: 'tiger', 10: 'lion', 11: 'jaguar', 12: 'bobcat', 13: 'monkey', 14: 'lemur', 15: 'tiger', 16: 'lion', 17: 'jaguar'}, 'Value': {0: 0, 1: 1, 2: 2, 3: 3, 4: 4, 5: 5, 6: 6, 7: 7, 8: 8, 9: 9, 10: 10, 11: 11, 12: 12, 13: 13, 14: 14, 15: 15, 16: 16, 17: 17}} g = pd.DataFrame(data).groupby(by=['buildings', 'vehicles', 'animals']).sum() # 核心处理 df = g.reset_index() parent_col = 'buildings' child_cols = [col for col in df.columns if col not in [parent_col, 'Value']] target_idx = set() for col in child_cols: # 统计每个值对应的不同父分组数量 val_parent_count = df.groupby(col)[parent_col].nunique() # 筛选跨父分组的重复值 duplicate_vals = val_parent_count[val_parent_count >= 2].index target_idx.update(df[df[col].isin(duplicate_vals)].index) # 恢复原多层索引格式 result = df.loc[list(target_idx)].set_index(g.index.names) print(result)
输出结果
Value buildings vehicles animals apartment big truck jaguar 11 lemur 8 lion 10 monkey 7 tiger 9 condo condo jaguar 17 lemur 14 lion 16 monkey 13 tiger 15
说明
代码自动适配子列数量,无需手动指定子列个数,最多支持10个及以上子列场景,仅需修改parent_col变量即可适配其他父分组列名的需求。
内容的提问来源于stack exchange,提问作者user2962397
相关产品推荐
相关产品推荐

