You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何优化处理10万+行DataFrame的嵌套For循环?

优化Pandas批量匹配竞品数据的性能方案

场景1:每行对应单个竞品索引

如果competitor_indexes是一维列表,每个元素对应df_competitor的单个行索引,直接用向量化索引就能彻底替代循环:

# 直接提取竞品数据的对应列,赋值给目标DataFrame
df['minimum_price'] = df_competitor['minimum_price'].loc[competitor_indexes].values
df['maximum_units'] = df_competitor['maximum_units'].loc[competitor_indexes].values

性能优势

Pandas的索引操作基于NumPy底层实现,是C级别的向量化运算,完全避开Python循环的解释执行开销,处理10万+行数据几乎瞬时完成。


场景2:每行对应多个竞品索引

如果competitor_indexes是嵌套列表(每个元素包含多个竞品行索引),需要计算对应组的最值,用分组聚合替代循环:

简洁高效实现

# 构造目标行索引与竞品索引的映射关系
comp_mapping = pd.Series(competitor_indexes, index=df.index).explode()

# 提取对应竞品数据并按目标行索引分组聚合
agg_results = df_competitor.loc[comp_mapping.values, ['minimum_price', 'maximum_units']].groupby(comp_mapping.index).agg(
    minimum_price=('minimum_price', 'min'),
    maximum_units=('maximum_units', 'max')
)

# 将聚合结果合并到目标DataFrame
df = df.join(agg_results)

核心逻辑

  1. explode()将嵌套列表展开,让每个竞品索引与对应的目标行索引一一绑定;
  2. 用loc批量提取竞品数据,再通过groupby按目标行索引分组,调用内置聚合函数计算最值;
  3. 最后用join合并结果,全程依赖Pandas的高效内置操作,无需Python级别的循环。

避坑提示

  • 提前校验competitor_indexes中的索引是否都存在于df_competitor中,可使用df_competitor.index.isin(...)排查,避免出现KeyError;
  • 若存在空的竞品索引组(元素为空列表),聚合后会生成NaN,可通过agg_results.fillna(...)设置默认值。

内容的提问来源于stack exchange,提问作者My name is jeff

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.16 07:05:52