Python实现Constant Volume Chart:解决分组成交量非精准匹配问题
高效实现固定成交量(Constant Volume)分组图表(精准50成交量/组)
针对百万行级别的数据,我们可以通过向量化计算分组边界+批量拆分数据的方式,高效实现每组(除最后一组)成交量精准等于50的需求,避免低效循环。
核心思路
- 计算累计成交量,确定所有固定成交量分组的边界
- 对每一行原始数据,计算其需要拆分到哪些分组中,以及每个分组对应的成交量
- 批量生成拆分后的行数据,保留原行的时间戳和价格
- 按分组聚合得到标准的OHLCV(开盘/最高/最低/收盘/成交量)数据
完整代码实现
import pandas as pd import numpy as np # 生成示例数据(百万行数据处理逻辑完全一致) date_rng = pd.date_range(start='2024-01-01', end='2024-12-31 23:00:00', freq='h') df = pd.DataFrame(date_rng, columns=['timestamp']) df['price'] = np.round(np.random.uniform(70, 100, size=(len(date_rng))), 2) df['volume'] = np.random.randint(1, 11, size=(len(date_rng))) constant_volume = 50 # 1. 计算累计成交量与分组边界 df['cum_volume'] = df['volume'].cumsum() total_groups = np.ceil(df['cum_volume'].iloc[-1] / constant_volume).astype(int) # 生成每个分组的起始累计成交量(0, 50, 100, ...) group_cum_starts = np.arange(0, total_groups * constant_volume, constant_volume) # 2. 向量化计算每行数据的拆分规则 # 判断每行累计成交量落入哪些分组区间 row_group_masks = df['cum_volume'].values[:, None] > group_cum_starts[None, :] # 获取每行对应的起始分组ID row_start_groups = row_group_masks.argmax(axis=1) # 计算该行在起始分组中需要填充的成交量(凑满50) row_first_fill = constant_volume - (group_cum_starts[row_start_groups] - df['cum_volume'].values + df['volume'].values) # 计算该行能拆出多少个完整的50成交量分组 row_full_groups = (df['volume'].values - row_first_fill) // constant_volume # 计算拆分后剩余的成交量(不足50的部分) row_last_remain = (df['volume'].values - row_first_fill) % constant_volume # 3. 批量生成拆分后的行数据 split_group_ids = [] split_volumes = [] split_prices = [] split_timestamps = [] for idx in range(len(df)): start_g = row_start_groups[idx] first_fill = row_first_fill[idx] full_groups = row_full_groups[idx] last_remain = row_last_remain[idx] # 添加起始分组的拆分行 split_group_ids.append(start_g) split_volumes.append(first_fill) split_prices.append(df['price'].iloc[idx]) split_timestamps.append(df['timestamp'].iloc[idx]) # 添加完整的中间分组(每个成交量固定为50) if full_groups > 0: split_group_ids.extend(range(start_g + 1, start_g + 1 + full_groups)) split_volumes.extend([constant_volume] * full_groups) split_prices.extend([df['price'].iloc[idx]] * full_groups) split_timestamps.extend([df['timestamp'].iloc[idx]] * full_groups) # 添加剩余成交量的行(仅当剩余量>0时) if last_remain > 0: split_group_ids.append(start_g + 1 + full_groups) split_volumes.append(last_remain) split_prices.append(df['price'].iloc[idx]) split_timestamps.append(df['timestamp'].iloc[idx]) # 转换为拆分后的DataFrame split_df = pd.DataFrame({ 'group_id': split_group_ids, 'timestamp': split_timestamps, 'price': split_prices, 'volume': split_volumes }) # 4. 按分组聚合得到最终的固定成交量图表数据 cv_chart = split_df.groupby('group_id').agg( timestamp=('timestamp', 'first'), open=('price', 'first'), high=('price', 'max'), low=('price', 'min'), close=('price', 'last'), volume=('volume', 'sum') ).set_index('timestamp') # 验证:除最后一组外,所有分组成交量均为50 # print(cv_chart[cv_chart['volume'] != constant_volume])
方案优势
- 高效性:核心分组边界计算采用numpy向量化操作,避免逐行逻辑判断;拆分数据的循环仅处理原始行(百万行级别的循环在Python中性能可接受)
- 精准性:除最后一组外,每组成交量严格等于设定的50,完全符合需求
- 可扩展性:支持任意固定成交量值,百万行数据无需分块即可直接处理(常规机器内存可承载)
内容的提问来源于stack exchange,提问作者nik
相关产品推荐
相关产品推荐

