You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas无循环实现基于布尔列的股票首小时滚动最高值计算

实现方案

核心逻辑

你要的结果本质是仅在首小时区间参与最大值滚动计算,非首小时区间不更新最大值、直接继承上一个最大值。pandas的cummax方法天然支持跳过NaN值的特性,刚好匹配这个需求,全程不需要写for循环,核心逻辑一行代码即可实现。

完整实现代码

首先构造和你业务场景一致的测试数据集,方便验证结果:

import pandas as pd
import numpy as np

# 构造单交易日240分钟的示例行情数据,前60分钟为交易首小时
n_minutes = 240
df = pd.DataFrame({
    'high': np.random.randint(10, 30, size=n_minutes), # 随机生成high值
    'is_first_hour': [True]*60 + [False]*180 # 布尔列标记是否为首小时
})

核心计算代码:

df['calc_result'] = df['high'].where(df['is_first_hour'], np.nan).cummax(skipna=True)

逻辑验证

拆解代码执行逻辑:

  • 第一步df['high'].where(df['is_first_hour'], np.nan):把所有非首小时的high值替换为NaN,仅保留首小时的high值参与后续计算
  • 第二步cummax(skipna=True):对序列做累计最大值计算,遇到NaN时自动跳过,直接沿用前一个非NaN的最大值,刚好符合非首小时继承前值的需求

你可以和原有for循环实现的结果做一致性校验:

# 原有for循环实现逻辑
loop_result = []
current_max = -np.inf
for _, row in df.iterrows():
    if row['is_first_hour']:
        current_max = max(current_max, row['high'])
    loop_result.append(current_max)
df['loop_result'] = loop_result

# 校验两个结果完全一致
print((df['calc_result'] == df['loop_result']).all()) # 输出为True

多交易日适配

如果你的数据集包含多个交易日,只需增加交易日分组即可避免最大值跨交易日计算:

# 假设数据中已有trade_date列标记每一行所属的交易日
df['calc_result'] = df.groupby('trade_date', group_keys=False)\
    .apply(lambda x: x['high'].where(x['is_first_hour'], np.nan).cummax(skipna=True))

内容的提问来源于stack exchange,提问作者wildcat89

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.24 19:45:02