You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何删除pandas DataFrame中符合Indent与MB列规则的多余行

样本数据代码

import pandas as pd
d = {'INDENT': {'0': 0, '1': 1, '2': 1, '3': 2, '4': 3, '5': 3, '6': 4, '7': 2, '8': 3}, 'MB': {'0': 'M', '1': 'B', '2': 'M', '3': 'B', '4': 'B', '5': 'M', '6': 'M', '7': 'B', '8': 'M'}}
df = pd.DataFrame(d)

原有嵌套循环代码

for index, row in df.iterrows():
    for row in range(index-1,0,-1):
        if df.loc[row].at["INDENT"] <= df.loc[index].at["INDENT"]-1:
            if df.loc[row].at["MB"]=="B":
                df.drop(df.index[index], inplace=True)
                break
            else:
                break

更新后样本数据

import pandas as pd
d = {
    'INDENT': {'0': 0, '1': 1, '2': 1, '3': 2, '4': 3, '5': 3, '6': 4, '7': 2, '8': 3}, 
    'MB': {'0': 'M', '1': 'B', '2': 'M', '3': 'B', '4': 'B', '5': 'M', '6': 'M', '7': 'B', '8': 'M'},
    'a': {'0': -1, '1': 5000, '2': 5000, '3': 5322, '4': 5449, '5': 5449, '6': 5621, '7': 5322, '8': 4666},
    'c': {'0': 5000, '1': 5222, '2': 5322, '3': 5449, '4': 5923, '5': 5621, '6': 5109, '7': 4666, '8': 5219}
    }
df = pd.DataFrame(d)

更新后图实现代码

import matplotlib.pyplot as plt
import networkx as nx
import pandas as pd
d = {
    'INDENT': {'0': 0, '1': 1, '2': 1, '3': 2, '4': 3, '5': 3, '6': 4, '7': 2, '8': 3}, 
    'MB': {'0': 'M', '1': 'B', '2': 'M', '3': 'B', '4': 'B', '5': 'M', '6': 'M', '7': 'B', '8': 'M'},
    'a': {'0': -1, '1': 5000, '2': 5000, '3': 5322, '4': 5449, '5': 5449, '6': 5621, '7': 5322, '8': 4666},
    'c': {'0': 5000, '1': 5222, '2': 5322, '3': 5449, '4': 5923, '5': 5621, '6': 5109, '7': 4666, '8': 5219}
    }
df = pd.DataFrame(d)

G = nx.Graph()
G = nx.from_pandas_edgelist(df, 'a', 'c', create_using=nx.DiGraph())
T = nx.dfs_tree(G, source=-1).reverse()
print([x for x in T])
nx.draw(G, with_labels=True)
plt.show()

解决方案

问题诊断

你原有嵌套循环失效的核心原因是遍历DataFrame时直接调用drop删除行,会导致索引实时错乱,后续遍历的下标和实际行无法对应。其次你写的内层循环遇到第一个父级INDENT就直接break,没有覆盖所有前置的B行规则,逻辑本身也不满足需求。

你尝试用networkx的思路走偏了,你用a和c列构建边和你的INDENT层级、MB标记规则完全无关,自然没法用来筛选待删除的行。如果没有额外的图分析需求,完全不需要引入networkx,直接线性遍历就能解决问题,效率更高逻辑也更清晰。

可直接运行的实现代码

我们用单遍历加阈值标记的逻辑实现需求,全程不修改原DataFrame的结构,避免索引错乱:

# 标记是否处于删除状态,以及对应的删除阈值
delete_threshold = None
keep_mask = []

for _, row in df.iterrows():
    current_indent = row['INDENT']
    current_mb = row['MB']
    # 先判断当前行是否需要删除
    if delete_threshold is not None:
        if current_indent > delete_threshold:
            keep_mask.append(False)
        else:
            # 遇到小于等于阈值的行,停止本次删除
            keep_mask.append(True)
            delete_threshold = None
    else:
        keep_mask.append(True)
    # 当前行保留且为B类时,设置后续的删除阈值
    if keep_mask[-1] and current_mb == 'B':
        delete_threshold = current_indent

# 过滤得到最终结果
df_result = df[keep_mask].reset_index(drop=True)

用你提供的样本数据运行,最终保留的行对应原索引为0、1、2、3、7,完全符合需求描述的规则。

内容的提问来源于stack exchange,提问作者onyex

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.04 21:45:01