如何检测DataFrame中价格穿越目标价并将对应Target置为NaN?
解决OHLC数据中目标价二次穿越的检测与清除问题
报错原因分析
你写的代码触发ValueError是因为:df['High'][i:] < df['Target'][i]返回的是一个布尔Series(一组布尔值),直接用if判断时,Python无法确定你要判断整个Series是全为真、任意为真还是其他条件,必须用.any()(任意元素为真)或.all()(所有元素为真)这类方法明确逻辑,但你的核心检测逻辑也需要调整,才能实现目标价二次穿越的检测。
实现思路
要完成需求,我们需要分步骤处理:
- 遍历所有包含有效Target值的行
- 对每个Target值T,先找到首次触及T的时间点(即Target行之后,第一个价格区间覆盖T的K线)
- 在首次触及之后的K线中,检测是否存在再次穿越T的情况(即价格从T的一侧移动到另一侧,或者K线区间直接覆盖T)
- 如果存在二次穿越,将原Target行的Target值设为NaN
代码实现
假设你的DataFrame索引是时间序列(Close Time),我们可以用向量化操作结合循环(针对非NaN的Target行,这类行数量通常远小于总数据量,效率可接受)来实现:
import pandas as pd # 复制原数据避免修改原始数据 df = df.copy() # 获取所有非NaN的Target行的索引和值 target_rows = df[df['Target'].notna()].index for idx in target_rows: t_value = df.loc[idx, 'Target'] # 获取当前Target行之后的所有K线,跳过当前行 future_data = df.loc[idx:, :].iloc[1:] if len(future_data) == 0: continue # 没有后续数据,直接跳过 # 1. 检测首次触及Target的位置:K线区间覆盖T(说明该时间段内价格触及过T) first_touch_mask = (future_data['High'] >= t_value) & (future_data['Low'] <= t_value) first_touch_idx = first_touch_mask.idxmax() if first_touch_mask.any() else None if first_touch_idx is None: continue # 从未触及过Target,跳过 # 2. 检测首次触及之后是否有二次穿越 post_touch_data = future_data.loc[first_touch_idx:, :].iloc[1:] if len(post_touch_data) == 0: continue # 二次穿越的两种判断场景: # a. K线区间直接覆盖T(波动过程中穿越) cross_mask = (post_touch_data['High'] >= t_value) & (post_touch_data['Low'] <= t_value) # b. 收盘价从T的一侧切换到另一侧(连续K线的穿越) prev_close = post_touch_data['Close'].shift(1) price_cross = ((prev_close > t_value) & (post_touch_data['Close'] < t_value)) | \ ((prev_close < t_value) & (post_touch_data['Close'] > t_value)) # 只要存在任意一种二次穿越情况,清除原Target值 if cross_mask.any() or price_cross.any(): df.loc[idx, 'Target'] = pd.NA
关键说明
- 优先用向量化操作处理批量数据,仅对非NaN的Target行循环,保证运行效率
- 用K线区间覆盖(
High >= T且Low <= T)判断“价格触及”,这是蜡烛图中最准确的判断方式,只要时间段内价格波动到过Target就算触及 - 二次穿越包含两种场景:单根K线直接覆盖Target,以及连续K线收盘价跨Target切换
内容的提问来源于stack exchange,提问作者Callum
相关产品推荐
相关产品推荐

