Pandas利用前序rank数值过滤当前行的实现方法
基于前序rank分组的数值过滤实现
过滤规则
若当前rank分组的某行,与所有前序rank分组中任意行相比,x、y的差值均在±1范围内,且z的差值在±0.1范围内,则剔除当前行。
初始数据示例
import pandas as pd df = pd.DataFrame({ 'rank': [1, 1, 2, 2, 3, 3], 'x': [0, 3, 0, 3, 4, 2], 'y': [0, 4, 0, 4, 5, 5], 'z': [1, 3, 1.2, 2.95, 3, 6], }) print(df) # 输出: # rank x y z # 0 1 0 0 1.00 # 1 1 3 4 3.00 # 2 2 0 0 1.20 # 3 2 3 4 2.95 # 4 3 4 5 3.00 # 5 3 2 5 6.00
实现代码
# 按rank升序排序,保证处理顺序正确 df_sorted = df.sort_values('rank', ignore_index=True) res = [] # 存储所有前序rank的保留行作为参考基准 ref_rows = [] for rank, group in df_sorted.groupby('rank'): # 最低rank的所有行直接保留 if not ref_rows: res.extend(group.to_dict('records')) ref_rows.extend(group[['x', 'y', 'z']].to_dict('records')) continue # 逐行判断当前rank的行是否需要剔除 for _, row in group.iterrows(): x, y, z = row['x'], row['y'], row['z'] drop_flag = False # 和所有前序参考行比对 for ref in ref_rows: if (abs(x - ref['x']) <= 1 and abs(y - ref['y']) <= 1 and abs(z - ref['z']) <= 0.1): drop_flag = True break if not drop_flag: res.append(row.to_dict()) # 保留的行加入参考基准,供后续更高rank比对 ref_rows.append({'x': x, 'y': y, 'z': z}) output = pd.DataFrame(res)
输出结果
print(output) # 输出: # rank x y z # 0 1 0 0 1.0 # 1 1 3 4 3.0 # 2 2 0 0 1.2 # 3 3 2 5 6.0
内容的提问来源于stack exchange,提问作者mike_gundy123
相关产品推荐
相关产品推荐

