如何让Pandas DataFrame中score1和score2列保持递增序列?
解决Pandas DataFrame移除历史中score1/score2高于当前行的行的问题
你的需求是遍历每行时,移除所有score1或score2值高于当前行的历史行,之前使用相邻行对比的掩码方法无法处理非相邻的历史行,因此达不到预期效果。
解决方案
通过遍历每行并维护一个有效行列表,动态过滤不符合条件的历史行:
import pandas as pd df = pd.DataFrame({ 'minutes':[1,2,3,4,5,6,7,8,9,10], 'score1': [0,0,1,1,0,0,0,1,1,2], 'score2': [0,1,1,1,1,2,1,1,2,2], 'sum_score': [0,1,2,2,1,2,1,2,3,4] }) valid_rows = [] for idx, row in df.iterrows(): # 过滤历史行中score1或score2高于当前行的记录 valid_rows = [r for r in valid_rows if not (r['score1'] > row['score1'] or r['score2'] > row['score2'])] valid_rows.append(row) # 转换为DataFrame并重置索引 result = pd.DataFrame(valid_rows).reset_index(drop=True) print(result)
输出结果
minutes score1 score2 sum_score 0 1 0 0 0 1 2 0 1 1 2 5 0 1 1 3 7 0 1 1 4 8 1 1 2 5 9 1 2 3 6 10 2 2 4
原理说明
该方法逐行处理数据:
- 对每个新行,先从已保留的有效列表中删除所有
score1大于当前行score1,或score2大于当前行score2的历史行 - 将当前行加入有效列表
- 最终将有效列表转换为DataFrame,得到符合需求的结果
内容的提问来源于stack exchange,提问作者mauriciokaminski
相关产品推荐
相关产品推荐

