You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中基于时间序列数据更新当日最高价列的问题

解决Pandas中计算累计最高价时出现0的问题

嘿,我帮你找出问题所在了!你遇到的0值问题,根源是对Pandas的向量化计算逻辑理解有点偏差,咱们一步步来解决:

问题原因

你的代码里用了np.where来判断更新high列,但**np.where是一次性完成所有行的向量化计算**,不会逐行实时更新high的值。你初始把high列全设为0,当条件不满足时,df.high.shift(1)调用的是原始未更新的high列的移位值(大部分是0),而不是前一行已经计算好的high值,这就导致不符合预期的0出现了。

解决方案

方案1:用cummax()(最简洁高效,推荐)

其实你要的就是从第一行到当前行的tradePrice累计最大值,Pandas已经内置了cummax()方法,直接就能实现需求,完全不需要手动写判断逻辑:

from matplotlib import style
import pandas as pd
import numpy as np
style.use('ggplot')
desired_width = 320
pd.set_option('display.width', desired_width)

df = pd.read_csv('C:/test.csv')
df.set_index('dateTime', inplace=True)
df.drop(['SecurityID'], axis=1, inplace=True)  # 新版本Pandas建议显式指定axis=1,避免警告

# 直接计算累计最高价
df['high'] = df['tradePrice'].cummax()

# 打印指定列查看结果
print(df[['tradePrice', 'high']].head(10))

运行后你会看到,每一行的high都会自动保留从第一行到当前行的最高tradePrice,完美符合你的需求:当前价格更高时更新,否则沿用之前的最高价。

方案2:逐行迭代(适合复杂自定义逻辑)

如果之后你需要更复杂的条件判断(比如结合其他列的规则),可以用逐行迭代的方式处理,虽然效率不如cummax(),但灵活性更高:

from matplotlib import style
import pandas as pd
import numpy as np
style.use('ggplot')
desired_width = 320
pd.set_option('display.width', desired_width)

df = pd.read_csv('C:/test.csv')
df.set_index('dateTime', inplace=True)
df.drop(['SecurityID'], axis=1, inplace=True)

# 初始化high列,第一行的high就是该行的tradePrice
df['high'] = df['tradePrice']

# 从第二行开始逐行判断更新
for idx in range(1, len(df)):
    current_price = df['tradePrice'].iloc[idx]
    prev_high = df['high'].iloc[idx-1]
    if current_price > prev_high:
        df['high'].iloc[idx] = current_price
    else:
        df['high'].iloc[idx] = prev_high

print(df[['tradePrice', 'high']].head(10))

再复盘下原来的代码问题

举个例子:第五行的tradePrice是12868,比前一行的high12870小,这时候np.where会取df.high.shift(1),但这个值是原始high列第四行的0(因为high列初始全是0,还没被逐行更新),而不是第四行已经计算好的12870,这就是为什么第五行的high变成了0。

内容的提问来源于stack exchange,提问作者Giladbi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 03:52:45