Python处理股票DataFrame:计算K线穿越期权行权价次数报错解决
问题:统计K线穿越期权行权价的次数并新增列
问题背景
我有包含股票OHLC(开盘价、最高价、最低价、收盘价)数据的DataFrame,想统计每一行对应的K线穿越期权行权价的次数,新增一列存储该统计值。
DataFrame示例
open high low close volume datetime datetime2 n_strike strk_diff pinned_min datetime2 2021-08-20 09:30:00-04:00 147.4400 147.5619 147.1201 147.3725 1660122.0 1629466200000 2021-08-20 13:30:00+00:00 145 2.3725 1 2021-08-20 09:31:00-04:00 147.3800 147.6600 147.1200 147.1350 430097.0 1629466260000 2021-08-20 13:31:00+00:00 145 2.1350 1 2021-08-20 09:32:00-04:00 147.1297 147.4800 147.0400 147.0550 308090.0 1629466320000 2021-08-20 13:32:00+00:00 145 2.0550 1 2021-08-20 09:33:00-04:00 147.1000 147.3199 147.0200 147.2348 285100.0 1629466380000 2021-08-20 13:33:00+00:00 145 2.2348 1 2021-08-20 09:34:00-04:00 147.2367 147.2600 146.9600 147.1250 290185.0 1629466440000 2021-08-20 13:34:00+00:00 145 2.1250 1 ... ... ... ... ... ... ... ... ... ... ... 2022-07-15 15:55:00-04:00 149.8900 149.9800 149.8400 149.9550 525630.0 1657914900000 2022-07-15 19:55:00+00:00 150 0.0450 0 2022-07-15 15:56:00-04:00 149.9600 150.0000 149.9100 149.9900 675573.0 1657914960000 2022-07-15 19:56:00+00:00 150 0.0100 0 2022-07-15 15:57:00-04:00 149.9900 150.0000 149.9400 149.9900 464692.0 1657915020000 2022-07-15 19:57:00+00:00 150 0.0100 0 2022-07-15 15:58:00-04:00 149.9900 150.0500 149.9200 150.0300 753358.0 1657915080000 2022-07-15 19:58:00+00:00 150 0.0300 0 2022-07-15 15:59:00-04:00 150.0300 150.2500 149.9700 150.1700 1978823.0 1657915140000 2022-07-15 19:59:00+00:00 150 0.1700
尝试的代码
#生成行权价列表 strikes = [*range(0,(round(df_expfri['high'].max())+5), 5)] for row in df_temp: H = df_temp['high'] L = df_temp['low'] count = 0 for x in strikes: if x < L : continue elif x > H: continue elif x > L & x < H: count +=1 print (count)
错误信息
--------------------------------------------------------------------------- ValueError Traceback (most recent call last) Input In [133], in <cell line: 7>() 10 count = 0 11 for x in strikes: ---> 12 if x < L : 13 continue 14 elif x > H: File C:\ProgramData\Anaconda3\lib\site-packages\pandas\core\generic.py:1535, in NDFrame.__nonzero__(self) 1533 @final 1534 def __nonzero__(self): -> 1535 raise ValueError( 1536 f"The truth value of a {type(self).__name__} is ambiguous. " 1537 "Use a.empty, a.bool(), a.item(), a.any() or a.all()." 1538 ) ValueError: The truth value of a Series is ambiguous. Use a.empty, a.bool(), a.item(), a.any() or a.all().
我猜测错误是因为H和L是Series类型,但不知道怎么解决,求帮助。
解决方案
错误原因
你的代码里,H = df_temp['high']和L = df_temp['low']取的是整个列的Series,而非每行的单个值,所以判断x < L时是拿单个数值和整个Series比较,Pandas无法确定布尔值逻辑,抛出歧义错误。另外,for row in df_temp默认遍历列名,不是行数据,循环逻辑完全错误。
方案1:逐行处理(逻辑直观,适合小数据量)
用apply逐行调用函数计算:
# 生成行权价列表 strikes = [*range(0, round(df_expfri['high'].max()) + 5, 5)] # 计算单条K线覆盖的行权价数量 def count_cross_strikes(row): low_val = row['low'] high_val = row['high'] # 统计落在[low, high]区间内的行权价数量 return sum(1 for strike in strikes if low_val <= strike <= high_val) # 新增列存储结果 df_temp['strike_cross_count'] = df_temp.apply(count_cross_strikes, axis=1)
方案2:向量化操作(效率极高,适合大数据量)
用NumPy广播实现批量比较,避免逐行循环:
import numpy as np # 转行权价为数组 strikes_arr = np.array(strikes) # 广播比较:每行的low/high和所有行权价做区间判断 in_range = (df_temp['low'].values[:, np.newaxis] <= strikes_arr) & (strikes_arr <= df_temp['high'].values[:, np.newaxis]) # 每行求和得到穿越次数 df_temp['strike_cross_count'] = in_range.sum(axis=1)
说明
两种方案都是统计行权价落在当前K线low到high区间内的数量,即K线穿越该行权价的次数(K线覆盖价格意味着存在上下穿越行为)。方案2的效率远高于方案1,数据量越大优势越明显。
内容的提问来源于stack exchange,提问作者J.Billman
相关产品推荐
相关产品推荐

