如何计算Tick数据分布中POC的上下一倍标准差?
问题背景
现有如下Tick数据:
,timestamp,close,security_code,volume,bid_volume,ask_volume 2024-04-02 01:00:00.128123+00:00,2024-04-02 01:00:00.128123+00:00,18465.5,NQ,1,0,1 2024-04-02 01:00:00.128123+00:00,2024-04-02 01:00:00.128123+00:00,18465.5,NQ,1,0,1 2024-04-02 01:00:03.782064+00:00,2024-04-02 01:00:03.782064+00:00,18465.25,NQ,1,0,1 2024-04-02 01:00:04.112603+00:00,2024-04-02 01:00:04.112603+00:00,18465.0,NQ,1,0,1 2024-04-02 01:00:04.112603+00:00,2024-04-02 01:00:04.112603+00:00,18465.0,NQ,1,0,1 2024-04-02 01:00:04.112603+00:00,2024-04-02 01:00:04.112603+00:00,18464.75,NQ,1,0,1 2024-04-02 01:00:04.112603+00:00,2024-04-02 01:00:04.112603+00:00,18464.75,NQ,1,0,1 2024-04-02 01:00:05.759876+00:00,2024-04-02 01:00:05.759876+00:00,18464.5,NQ,1,0,1 2024-04-02 01:00:06.273686+00:00,2024-04-02 01:00:06.273686+00:00,18464.75,NQ,5,5,0
已通过以下Python代码计算出实时最高价(high)、最低价(low)和成交量密集点(POC):
import pandas as pd, matplotlib.pyplot as plt from collections import defaultdict df = pd.read_csv("csv/nq_out_daily.csv") df.drop('Unnamed: 0', inplace=True, axis=1) df['timestamp'] = pd.to_datetime(df['timestamp']) df["timestamp"] = df['timestamp'].dt.strftime('%d-%m-%Y %H:%M:%S') summary = {"high": [], "low": [], "poc": []} dist = defaultdict(float) current_high = current_low = None for idx, (timestamp, tick, ask, bid) in enumerate(zip(df.timestamp, df.close, df.ask_volume, df.bid_volume)): current_high = tick if (current_high is None or tick > current_high) else current_high current_low = tick if (current_low is None or tick < current_low) else current_low dist[tick] += 1 summary["high"].append(current_high) summary["low"].append(current_low) summary["poc"].append(max(dist, key=dist.get)) # plot the summary fig = plt.figure() x = range(len(summary["high"])) plt.scatter(x, summary["high"], s=1) plt.scatter(x, summary["low"], s=1) plt.scatter(x, summary["poc"], s=1) plt.legend(['high', 'low', 'poc']) plt.savefig(f"distribution.png") plt.close(fig)
问题
如何计算POC的上下一倍标准差?
解决方案
要计算POC的上下一倍标准差,需基于价格对应的成交量分布计算加权标准差(权重为各价格的累计成交量),再用POC值加减该标准差得到区间。具体步骤和修改后的代码如下:
核心逻辑
- 利用已有的
dist字典维护每个价格的累计成交量 - 计算价格的加权平均值(以成交量为权重)
- 计算加权标准差:公式为 $\sigma = \sqrt{\frac{\sum w_i \times (x_i - \mu)^2}{\sum w_i}}$,其中$\mu$是加权均值,$w_i$是价格$x_i$的累计成交量
- 实时计算POC的上下一倍标准差并加入结果集合
修改后的完整代码
import pandas as pd, matplotlib.pyplot as plt from collections import defaultdict import math df = pd.read_csv("csv/nq_out_daily.csv") df.drop('Unnamed: 0', inplace=True, axis=1) df['timestamp'] = pd.to_datetime(df['timestamp']) df["timestamp"] = df['timestamp'].dt.strftime('%d-%m-%Y %H:%M:%S') summary = {"high": [], "low": [], "poc": [], "poc_upper_std": [], "poc_lower_std": []} dist = defaultdict(float) current_high = current_low = None for idx, (timestamp, tick, ask, bid) in enumerate(zip(df.timestamp, df.close, df.ask_volume, df.bid_volume)): # 更新实时高低点和成交量分布 current_high = tick if (current_high is None or tick > current_high) else current_high current_low = tick if (current_low is None or tick < current_low) else current_low dist[tick] += 1 # 获取当前POC current_poc = max(dist, key=dist.get) # 计算加权均值和加权标准差 total_volume = sum(dist.values()) weighted_mean = sum(price * vol for price, vol in dist.items()) / total_volume weighted_variance = sum(vol * (price - weighted_mean)**2 for price, vol in dist.items()) / total_volume weighted_std = math.sqrt(weighted_variance) # 计算POC上下一倍标准差 poc_upper = current_poc + weighted_std poc_lower = current_poc - weighted_std # 存入结果集合 summary["high"].append(current_high) summary["low"].append(current_low) summary["poc"].append(current_poc) summary["poc_upper_std"].append(poc_upper) summary["poc_lower_std"].append(poc_lower) # 新增标准差曲线的绘图 fig = plt.figure() x = range(len(summary["high"])) plt.scatter(x, summary["high"], s=1) plt.scatter(x, summary["low"], s=1) plt.scatter(x, summary["poc"], s=1) plt.scatter(x, summary["poc_upper_std"], s=1, color='orange') plt.scatter(x, summary["poc_lower_std"], s=1, color='orange') plt.legend(['high', 'low', 'poc', 'poc_upper_std', 'poc_lower_std']) plt.savefig(f"distribution_with_std.png") plt.close(fig)
说明
- 代码中每次循环都会实时计算截至当前tick的加权标准差,确保
poc_upper_std和poc_lower_std是随时间更新的实时值 - 加权标准差考虑了不同价格的成交量权重,更贴合成交量分布的实际离散程度
- 新增的绘图代码会把POC的上下标准差曲线也画出来,方便可视化观察
内容的提问来源于stack exchange,提问作者Jan
相关产品推荐
相关产品推荐

