You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何迭代pandas DataFrame的列并按条件删除对应行(Jupyter Notebook)

解决方法

推荐方案:用pandas向量化操作(无需遍历,效率更高)

你的需求本质是筛选出Coaches列与上一行或下一行值不同的行,直接用shift方法偏移对比即可:

import pandas as pd
df = pd.read_csv("hawks.csv")

# 情况1:按Coaches列完整字符串对比
cond = (df['Coaches'] != df['Coaches'].shift(1)) | (df['Coaches'] != df['Coaches'].shift(-1))

# 情况2:仅对比教练姓名,忽略后面的战绩(匹配你举的示例中M. Budenholzer几行判定为相同的规则)
# 先提取姓名部分
# df['coach_name'] = df['Coaches'].str.split('(').str[0].str.strip()
# cond = (df['coach_name'] != df['coach_name'].shift(1)) | (df['coach_name'] != df['coach_name'].shift(-1))

# 筛选符合条件的行
res_df = df[cond]

该方法自动适配首尾行的边界情况,首行默认和不存在的上一行判定为不同,末行默认和不存在的下一行判定为不同,符合你的筛选规则。

遍历实现方案(按你需要的迭代逻辑实现)

如果一定要用逐行迭代的方式实现,可以参考以下代码:

import pandas as pd
df = pd.read_csv("hawks.csv")

keep_idx = []
row_count = len(df)
for i in range(row_count):
    current_val = df.loc[i, 'Coaches']
    # 首行只需对比下一行
    if i == 0:
        if current_val != df.loc[i+1, 'Coaches']:
            keep_idx.append(i)
    # 末行只需对比上一行
    elif i == row_count - 1:
        if current_val != df.loc[i-1, 'Coaches']:
            keep_idx.append(i)
    # 中间行对比上下两行
    else:
        prev_val = df.loc[i-1, 'Coaches']
        next_val = df.loc[i+1, 'Coaches']
        if current_val != prev_val or current_val != next_val:
            keep_idx.append(i)

res_df = df.loc[keep_idx]

注意:尽量不要用iloc按索引位置取列,直接通过列名df['Coaches']取值更稳妥,不会因为csv列顺序变化导致取值错误。

内容的提问来源于stack exchange,提问作者prismarine

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.02 01:36:02