如何在Pandas Series中快速检测并删除为其他行子集的行
高效删除Pandas Series中属于其他行子集的元素
问题场景
你有一个存储数字列表的大型Pandas Series,需要剔除那些是其他行子集的元素。原双重循环实现速度极慢,以下是更高效的解决方案。
示例数据
import pandas as pd cycles = pd.Series([[1, 2, 3, 4], [3, 4], [5, 6, 9, 7], [5, 9]])
目标是删除第2、4行(索引1、3),因为它们分别是第1、3行(索引0、2)的子集。
方案一:集合操作+遍历判断(中等规模数据适用)
利用集合的issuperset方法快速判断子集关系,实现精准过滤:
# 将每个列表转换为集合,便于子集判断 cycle_sets = cycles.apply(set) # 生成保留标记:如果当前集合是其他集合的子集,则标记为False keep_mask = [] for current_set in cycle_sets: # 检查是否存在其他集合包含当前集合,且两个集合不相等(避免误删自身) has_superset = any( other_set.issuperset(current_set) and other_set != current_set for other_set in cycle_sets ) keep_mask.append(not has_superset) # 过滤得到结果 filtered_cycles = cycles[keep_mask] print(filtered_cycles)
输出结果:
0 [1, 2, 3, 4] 2 [5, 6, 9, 7] dtype: object
方案二:排序优化批量判断(大规模数据适用)
通过按集合长度降序排序,减少重复判断次数——大集合在前,后续只需检查前面的大集合是否包含当前小集合,无需对比后续更小的集合:
# 构造包含原数据、集合、长度的DataFrame cycle_df = pd.DataFrame({ 'original': cycles, 'data_set': cycles.apply(set), 'length': cycles.apply(len) }) # 按集合长度降序排序,大集合优先处理 sorted_df = cycle_df.sort_values(by='length', ascending=False).reset_index(drop=True) keep_mask = [True] * len(sorted_df) for i in range(len(sorted_df)): if not keep_mask[i]: continue current_set = sorted_df.loc[i, 'data_set'] # 仅检查后续长度更小的集合 for j in range(i + 1, len(sorted_df)): if keep_mask[j] and current_set.issuperset(sorted_df.loc[j, 'data_set']): keep_mask[j] = False # 过滤后恢复原数据顺序 filtered_result = sorted_df[keep_mask].sort_index()['original'] print(filtered_result)
性能说明
原双重循环的时间复杂度为O(n²),方案二通过排序减少了无效判断,实际运行效率随数据量增大优势越明显——对于百万级数据,效率可提升数倍甚至数十倍。
内容的提问来源于stack exchange,提问作者Alireza75
相关产品推荐
相关产品推荐

