为何df.drop删除行数超出预期?英超DataFrame处理异常
问题分析与解决
可能的原因及修复方案
1. 未保存删除操作的结果
pandas的drop()方法默认返回新的DataFrame,不会修改原数据。如果只执行other_team_df.drop(del_row_index)却没有将结果赋值给变量,也未添加inplace=True参数,原other_team_df不会发生任何变化。你看到的72行大概率是误查看了其他数据,或是之前操作残留的错误结果。
修复方式二选一:
- 重新赋值给原变量或新变量:
other_team_df = other_team_df.drop(del_row_index) - 使用
inplace=True直接修改原DataFrame:other_team_df.drop(del_row_index, inplace=True)
2. 重复索引导致误删
如果你的DataFrame存在重复的行索引,drop()会删除所有匹配该索引值的行,而非仅你筛选出的1596行。比如某索引值在del_row_index中出现1次,但原DataFrame里有几十行共用该索引,这些行都会被批量删除,最终剩余行数远低于预期。
修复步骤:
- 先检查索引是否唯一:
print(other_team_df.index.is_unique) - 若返回
False,先重置索引再执行删除:# 重置索引,丢弃原重复索引 other_team_df = other_team_df.reset_index(drop=True) # 重新生成待删除索引(用isin简化条件) target_teams = ['Arsenal', 'Chelsea', 'Liverpool', 'Tottenham', 'Man City', 'Man United'] del_row_index = other_team_df[other_team_df['HomeTeam'].isin(target_teams)].index # 执行删除 other_team_df = other_team_df.drop(del_row_index)
优化建议
用isin()替代多个|逻辑,让代码更简洁易维护,就是上面示例里的写法。
内容的提问来源于stack exchange,提问作者Immanuel Nii Odai Odarteifio
相关产品推荐
相关产品推荐

