You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas分组计算股价移动平均及占比:求高效优化实现方案

高效计算个股移动平均及股价站线占比的Pandas优化方案

优化后代码

# 1. 计算20日滚动均线,自动对齐原DataFrame
prices_df['sma_20'] = prices_df.groupby('ticker')['closeadj'].transform(lambda x: x.rolling(20).mean())

# 2. 标记股价是否高于均线,布尔值转int更高效
prices_df['above_sma_20'] = (prices_df['closeadj'] > prices_df['sma_20']).astype(int)

# 3. 直接计算滚动252天的站线占比,省去额外除法步骤
prices_df['above_sma_20_pct'] = prices_df.groupby('ticker')['above_sma_20'].transform(lambda x: x.rolling(252).mean())

核心优化点

  • 用transform替代手动索引对齐:原代码中groupby.rolling().reset_index(0, drop=True)需要手动处理索引,transform会自动返回与原DataFrame行对齐的结果,避免索引操作的额外开销,代码更简洁。
  • 布尔值转int替代np.where:(closeadj > sma_20)直接生成布尔Series,调用.astype(int)比np.where更高效,Pandas对布尔类型的操作有底层优化。
  • 滚动均值直接计算占比:原代码先求和再除以252,等价于直接对above_sma_20做滚动均值,减少一次中间列的生成与计算,节省内存和时间。

进阶优化(超大数据集场景)

如果个股数量多、数据量极大,可额外做以下优化:

  • 将ticker转为分类类型:prices_df['ticker'] = prices_df['ticker'].astype('category'),Pandas对分类类型的分组操作效率远高于字符串类型。
  • 合并步骤减少中间列(若无需保留above_sma_20):
prices_df['sma_20'] = prices_df.groupby('ticker')['closeadj'].transform(lambda x: x.rolling(20).mean())
prices_df['above_sma_20_pct'] = prices_df.groupby('ticker').apply(
    lambda g: (g['closeadj'] > g['sma_20']).rolling(252).mean()
).reset_index(0, drop=True)

内容的提问来源于stack exchange,提问作者ng150716

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.09 17:25:26