You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas:无循环筛选'hbr'列值大于后续n行的行

Pandas无循环筛选:hbr值大于后续n行所有对应值的行

问题背景

给定如下Pandas DataFrame:

index       bsl     mb              hbr
74729   47091.89    47160.00        74731.0
74730   47159.00    47201.00        74835.0
74731   47196.50    47201.50        74735.0
74732   47186.50    47198.02        74735.0
74733   47191.50    47191.50        74735.0
74734   47162.50    47254.00        74736.0
74735   47252.50    47411.50        74736.0
74736   47414.50    47421.00        74747.0
74737   47368.50    47403.00        74742.0
74738   47305.00    47310.00        74742.0
74739   47292.00    47320.00        74742.0
74740   47302.00    47374.00        74742.0
74741   47291.47    47442.50        74899.0
74742   47403.50    47416.50        74746.0
74743   47354.34    47362.50        74746.0
74737   47368.50    47403.00        74742.0

需要无循环实现:筛选出所有hbr列值大于后续n行所有对应hbr值的行(n可大于500)。以n=3为例,预期输出为:

index       bsl     mb              hbr
74730   47159.00    47201.00        74835.0
74736   47414.50    47421.00        74747.0
74741   47291.47    47442.50        74899.0

解决方案

利用Pandas的滚动窗口函数实现高效无循环筛选,核心思路是计算每一行后续n行的hbr最大值,再与当前行hbr比较:

import pandas as pd

# 构造原始DataFrame(如果已有数据可跳过此步)
data = {
    'index': [74729,74730,74731,74732,74733,74734,74735,74736,74737,74738,74739,74740,74741,74742,74743,74737],
    'bsl': [47091.89,47159.00,47196.50,47186.50,47191.50,47162.50,47252.50,47414.50,47368.50,47305.00,47292.00,47302.00,47291.47,47403.50,47354.34,47368.50],
    'mb': [47160.00,47201.00,47201.50,47198.02,47191.50,47254.00,47411.50,47421.00,47403.00,47310.00,47320.00,47374.00,47442.50,47416.50,47362.50,47403.00],
    'hbr': [74731.0,74835.0,74735.0,74735.0,74735.0,74736.0,74736.0,74747.0,74742.0,74742.0,74742.0,74742.0,74899.0,74746.0,74746.0,74742.0]
}
df = pd.DataFrame(data).set_index('index')

n = 3  # 可替换为任意大于0的整数

# 计算当前行后续n行的hbr最大值,最后n行因无足够后续行返回NaN
next_n_hbr_max = df['hbr'].shift(-1).rolling(window=n).max()

# 构造筛选条件:当前hbr > 后续n行最大值,且排除无足够后续行的最后n行
filter_mask = df['hbr'] > next_n_hbr_max
filter_mask = filter_mask.dropna()

# 筛选结果
result_df = df[filter_mask]

print(result_df)

代码解释

  • df['hbr'].shift(-1):将hbr列整体上移一行,使第i行位置对应原数据第i+1行的hbr值,为后续滚动计算做准备
  • rolling(window=n).max():对移位后的hbr列取窗口大小为n的滚动最大值,得到当前行之后连续n行的hbr最大值
  • filter_mask.dropna():过滤掉最后n行(这些行没有足够的后续n行,滚动计算结果为NaN)
  • 布尔索引df[filter_mask]:直接筛选出符合条件的行,全程无循环,效率极高,适合n大于500的大场景

内容的提问来源于stack exchange,提问作者Sina

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.23 13:10:29