如何删除pandas DataFrame中符合Indent与MB列规则的多余行
样本数据代码
import pandas as pd d = {'INDENT': {'0': 0, '1': 1, '2': 1, '3': 2, '4': 3, '5': 3, '6': 4, '7': 2, '8': 3}, 'MB': {'0': 'M', '1': 'B', '2': 'M', '3': 'B', '4': 'B', '5': 'M', '6': 'M', '7': 'B', '8': 'M'}} df = pd.DataFrame(d)
原有嵌套循环代码
for index, row in df.iterrows(): for row in range(index-1,0,-1): if df.loc[row].at["INDENT"] <= df.loc[index].at["INDENT"]-1: if df.loc[row].at["MB"]=="B": df.drop(df.index[index], inplace=True) break else: break
更新后样本数据
import pandas as pd d = { 'INDENT': {'0': 0, '1': 1, '2': 1, '3': 2, '4': 3, '5': 3, '6': 4, '7': 2, '8': 3}, 'MB': {'0': 'M', '1': 'B', '2': 'M', '3': 'B', '4': 'B', '5': 'M', '6': 'M', '7': 'B', '8': 'M'}, 'a': {'0': -1, '1': 5000, '2': 5000, '3': 5322, '4': 5449, '5': 5449, '6': 5621, '7': 5322, '8': 4666}, 'c': {'0': 5000, '1': 5222, '2': 5322, '3': 5449, '4': 5923, '5': 5621, '6': 5109, '7': 4666, '8': 5219} } df = pd.DataFrame(d)
更新后图实现代码
import matplotlib.pyplot as plt import networkx as nx import pandas as pd d = { 'INDENT': {'0': 0, '1': 1, '2': 1, '3': 2, '4': 3, '5': 3, '6': 4, '7': 2, '8': 3}, 'MB': {'0': 'M', '1': 'B', '2': 'M', '3': 'B', '4': 'B', '5': 'M', '6': 'M', '7': 'B', '8': 'M'}, 'a': {'0': -1, '1': 5000, '2': 5000, '3': 5322, '4': 5449, '5': 5449, '6': 5621, '7': 5322, '8': 4666}, 'c': {'0': 5000, '1': 5222, '2': 5322, '3': 5449, '4': 5923, '5': 5621, '6': 5109, '7': 4666, '8': 5219} } df = pd.DataFrame(d) G = nx.Graph() G = nx.from_pandas_edgelist(df, 'a', 'c', create_using=nx.DiGraph()) T = nx.dfs_tree(G, source=-1).reverse() print([x for x in T]) nx.draw(G, with_labels=True) plt.show()
解决方案
问题诊断
你原有嵌套循环失效的核心原因是遍历DataFrame时直接调用drop删除行,会导致索引实时错乱,后续遍历的下标和实际行无法对应。其次你写的内层循环遇到第一个父级INDENT就直接break,没有覆盖所有前置的B行规则,逻辑本身也不满足需求。
你尝试用networkx的思路走偏了,你用a和c列构建边和你的INDENT层级、MB标记规则完全无关,自然没法用来筛选待删除的行。如果没有额外的图分析需求,完全不需要引入networkx,直接线性遍历就能解决问题,效率更高逻辑也更清晰。
可直接运行的实现代码
我们用单遍历加阈值标记的逻辑实现需求,全程不修改原DataFrame的结构,避免索引错乱:
# 标记是否处于删除状态,以及对应的删除阈值 delete_threshold = None keep_mask = [] for _, row in df.iterrows(): current_indent = row['INDENT'] current_mb = row['MB'] # 先判断当前行是否需要删除 if delete_threshold is not None: if current_indent > delete_threshold: keep_mask.append(False) else: # 遇到小于等于阈值的行,停止本次删除 keep_mask.append(True) delete_threshold = None else: keep_mask.append(True) # 当前行保留且为B类时,设置后续的删除阈值 if keep_mask[-1] and current_mb == 'B': delete_threshold = current_indent # 过滤得到最终结果 df_result = df[keep_mask].reset_index(drop=True)
用你提供的样本数据运行,最终保留的行对应原索引为0、1、2、3、7,完全符合需求描述的规则。
内容的提问来源于stack exchange,提问作者onyex
相关产品推荐
相关产品推荐

