You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Pandas高效生成股票OHLC数据的三级嵌套字典结构

高效实现10万条DataFrame转三级嵌套OHLC字典(支持实时更新)

核心高效转换代码

假设你的DataFrame包含列:stock_code(股票代码)、datetime(日期时间)、open、high、low、close,用Pandas内置的分组+向量化操作实现转换,避免手动循环的高耗时:

import pandas as pd

# 核心转换逻辑
ohlc_dict = df.groupby('stock_code').apply(
    lambda group: group.set_index('datetime')
                      .apply(lambda row: row.tolist(), axis=1)
                      .to_dict()
).to_dict()

转换后结构完全符合需求:

{
    "股票代码A": {
        时间戳1: [开盘价, 最高价, 最低价, 收盘价],
        时间戳2: [开盘价, 最高价, 最低价, 收盘价],
        ...
    },
    "股票代码B": {
        ...
    }
}

为什么这个方案高效?

  • 全程用Pandas的C底层优化操作,避免Python层面的逐行循环,10万条数据转换耗时通常在2-5秒内
  • groupby直接按股票代码分组,set_index('datetime')将时间作为二级字典的键,apply(row.tolist())快速生成OHLC列表

实时更新优化方案

针对实时场景的更新需求,直接操作字典而非重新转换整个DataFrame,实现O(1)时间复杂度的更新:

def update_single_ohlc(ohlc_dict, stock_code, datetime_key, new_ohlc):
    # 若股票不存在则初始化空字典
    if stock_code not in ohlc_dict:
        ohlc_dict[stock_code] = {}
    # 直接覆盖或新增时间对应的OHLC值
    ohlc_dict[stock_code][datetime_key] = new_ohlc

# 调用示例
update_single_ohlc(
    ohlc_dict, 
    "000001", 
    pd.Timestamp("2024-01-01 09:30:00"), 
    [12.5, 12.8, 12.3, 12.6]
)

常见问题排查

如果之前转换结构异常,大概率是这两个原因:

  • datetime列不是可哈希类型(比如未转成datetime64[ns]或字符串),无法作为字典键,转换前可执行df['datetime'] = pd.to_datetime(df['datetime'])
  • 手动循环时重复遍历股票代码,导致逻辑混乱且效率低下,直接用groupby可避免这类问题

内容的提问来源于stack exchange,提问作者lightspeed192

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.04 13:15:31