You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将嵌套循环Python代码转为Pandas/NumPy Vectorization实现以提速

优化嵌套循环的Pandas向量化方案

原代码的嵌套for循环逐行操作DataFrame是性能瓶颈核心——Pandas对单元素操作的开销极大,数据量上去后耗时会爆炸。下面是用向量化操作重构的版本,把内层循环替换成批量处理,同时完全保留原逻辑的正确性:

import pandas as pd
import numpy as np

balance = 10000

raw_data = [[1,2,4,1,3],[2,3,7,2,4],[3,4,5,3,4],[4,4,9,1,5],[5,5,6,4,5]]
raw_df = pd.DataFrame(raw_data, columns=['D','O','H','L','C'])

history_data = [[1,1,5,np.nan,4],[0,1,3,np.nan,4],[1,0,4,2,3],[1,0,1,6,0],[0,1,7,np.nan,8]]
history_df = pd.DataFrame(history_data, columns=['TY','ST','OP','CL','SL'])

# 按顺序处理raw_df的每一行(因为新增行要参与后续循环,外层只能逐行执行,但内层用批量操作)
for _, row in raw_df.iterrows():
    current_L = row['L']
    # 一次性筛选出所有符合条件的行:ST=1、TY=1、SL>=当前L
    match_mask = (history_df['ST'] == 1) & (history_df['TY'] == 1) & (history_df['SL'] >= current_L)
    # 统计匹配行数,直接更新balance
    balance += match_mask.sum() * 20
    # 批量更新符合条件的行的CL和ST
    history_df.loc[match_mask, 'CL'] = current_L
    history_df.loc[match_mask, 'ST'] = 0
    
    # 处理新增行逻辑,用concat替代已弃用的append
    if row['C'] > 4:
        new_row = pd.DataFrame({'TY':[0], 'ST':[1], 'OP':[5], 'CL':[np.nan], 'SL':[9]})
        history_df = pd.concat([history_df, new_row], ignore_index=True)

核心优化点:

  • 内层循环替换为布尔掩码批量筛选:用Pandas的布尔索引一次性定位所有符合条件的行,底层是C级别的向量化操作,比Python循环快几个数量级。
  • 批量赋值替代逐行修改:通过loc[mask, col]一次性更新所有匹配行,避免了逐行操作的巨大开销。
  • 用concat替代append:append已经被官方弃用,pd.concat在多次新增行时效率更高。
  • 直接统计匹配数:match_mask.sum()直接计算布尔数组中True的数量,比循环计数高效得多。

如果raw_df的行数特别多,外层的iterrows还可以进一步优化,但因为每次迭代会新增行到history_df,必须按顺序处理每一行,所以这个外层循环是必要的——但内层的批量操作已经解决了90%以上的性能问题。

内容的提问来源于stack exchange,提问作者576KB

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.07 08:10:37