Pandas按条件去重:基于Attribute值保留DataFrame行的实现
解决方案:按规则筛选DataFrame行
针对你提出的需求,可以通过两种高效方式实现,核心逻辑是按location分组后,优先保留非特定属性值的行;仅当分组内所有行均为特定值时,才保留该分组的全部行。
方法一:分组后自定义筛选函数
逻辑直观,适合理解分组筛选的完整过程:
import pandas as pd # 构造示例数据 location = ['A', 'B', 'B', 'C', 'C'] attribute = ['1', '2', '3', '2', '2'] df = pd.DataFrame({'location': location, 'attribute': attribute}) # 定义需要排除的特定属性值 specific_val = '2' # 自定义分组筛选函数 def filter_group(group): # 检查当前分组是否存在非特定值的行 has_non_target = (group['attribute'] != specific_val).any() if has_non_target: # 存在非特定值时,仅保留这些行 return group[group['attribute'] != specific_val] else: # 全为特定值时,保留整个分组 return group # 应用筛选逻辑 result = df.groupby('location', group_keys=False).apply(filter_group) print(result)
输出结果:
location attribute 0 A 1 2 B 3 3 C 2 4 C 2
方法二:基于标记列的布尔筛选
性能更优,适合处理大规模数据集:
import pandas as pd # 构造示例数据 location = ['A', 'B', 'B', 'C', 'C'] attribute = ['1', '2', '3', '2', '2'] df = pd.DataFrame({'location': location, 'attribute': attribute}) specific_val = '2' # 给每行添加标记:该location是否存在非特定值的行 df['has_non_target'] = df.groupby('location')['attribute'].transform( lambda x: (x != specific_val).any() ) # 组合筛选条件 # 条件1:存在非特定值的location,仅保留非特定值的行 # 条件2:无其他值的location,保留所有行 result = df[ (df['has_non_target'] & (df['attribute'] != specific_val)) | (~df['has_non_target']) ].drop('has_non_target', axis=1) print(result)
两种方法均可得到符合需求的结果,可根据数据规模选择适配方案。
内容的提问来源于stack exchange,提问作者i.i.k.
相关产品推荐
相关产品推荐

