You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何修复单独运行快速但批量循环时耗时极长的for循环问题?

批量循环处理日期子集时陷入“无限运行”的问题排查与优化方案

问题背景

你提到单独处理单个日期子集时代码运行很快,但批量循环60次就好像陷入无限运行,核心需求是给每个日期子集生成col3后合并成完整DataFrame。先看一下你的原始代码:

import pandas as pd
import numpy as np

def reweight(weights, cap):
    # Obtain constrained weights
    constrained_wts = np.minimum(cap, weights)
    # Locate all stocks with less than max weight
    nonmax = constrained_wts.ne(cap)
    # Calculate adjustment factor - this is proportional to original weights
    adj = ((1 - constrained_wts.sum()) * weights.loc[nonmax] / weights.loc[nonmax].sum())
    # Apply adjustment to obtain final weights
    constrained_wts = constrained_wts.mask(nonmax, weights + adj)
    # Repeat process in loop till conditions are satisfied
    while ((constrained_wts.sum() < 1) or (len(constrained_wts[constrained_wts > cap]) >=1 )):
        # Obtain constrained weights
        constrained_wts = np.minimum(cap, constrained_wts)
        # Locate all stocks with less than max weight
        nonmax = constrained_wts.ne(cap)
        # Calculate adjustment factor - this is proportional to original weights
        adj = ((1 - constrained_wts.sum()) * constrained_wts.loc[nonmax] / weights.loc[nonmax].sum())
        # Apply adjustment to obtain final weights
        constrained_wts = constrained_wts.mask(nonmax, constrained_wts + adj)
    return constrained_wts

# 原始主代码
data = pd.read_csv("df.csv")
data.YYMM = data.YYMM.apply(pd.to_datetime)
dates = data.groupby(data.YYMM).sum().index.values
data1 = pd.DataFrame()
for i in dates:
    df1 = data[data.YYMM == i]
    df1 = df1.sort_values(by='col1', ascending=False)
    df1['col2'] = df1.col1 / sum(df1.col1)
    df1['col3'] = reweight(df1.col2, cap)
    data1 = data1.append(df1, ignore_index = True)

核心问题分析

  • DataFrame.append的低效性:
    在循环里反复调用data1.append(df1)是最大的性能杀手。每次append都会创建一个新的DataFrame对象,随着data1的体积越来越大,内存复制的开销会呈指数级增长,60次循环下来会变得异常缓慢,看起来像是“无限运行”。
  • reweight函数的while循环可能死循环:
    你的while循环条件依赖浮点数的精确判断,加上部分日期子集的权重调整逻辑可能无法收敛,会导致循环一直跑下去,直接卡住程序。
  • 日期子集切片的低效性:
    用data[data.YYMM == i]在循环里反复筛选数据,不如直接用pandas.groupby的apply方法,它能更高效地按日期分组处理。

优化后的解决方案

方案1:修复append低效+避免死循环

把每个处理后的日期子集存入列表,最后一次性用pd.concat合并,同时给reweight函数加循环次数限制和浮点数误差容忍:

import pandas as pd
import numpy as np

def reweight(weights, cap):
    constrained_wts = np.minimum(cap, weights)
    nonmax = constrained_wts.ne(cap)
    
    # 处理极端情况:所有权重都等于cap,直接返回
    if nonmax.sum() == 0:
        return constrained_wts
    
    adj = ((1 - constrained_wts.sum()) * weights.loc[nonmax] / weights.loc[nonmax].sum())
    constrained_wts = constrained_wts.mask(nonmax, weights + adj)
    
    # 加最大迭代次数+浮点数误差容忍,防止死循环
    max_iter = 1000
    iter_count = 0
    # 用1e-8处理浮点数精度问题,避免因微小误差导致循环无法终止
    while ((constrained_wts.sum() < 1 - 1e-8) or (constrained_wts.gt(cap + 1e-8).any())) and iter_count < max_iter:
        constrained_wts = np.minimum(cap, constrained_wts)
        nonmax = constrained_wts.ne(cap)
        if nonmax.sum() == 0:
            break
        adj = ((1 - constrained_wts.sum()) * constrained_wts.loc[nonmax] / weights.loc[nonmax].sum())
        constrained_wts = constrained_wts.mask(nonmax, constrained_wts + adj)
        iter_count += 1
    
    # 若达到最大迭代次数,打印警告便于排查
    if iter_count >= max_iter:
        print(f"Warning: Reweight did not converge after {max_iter} iterations")
    return constrained_wts

# 优化后的主代码
data = pd.read_csv("df.csv")
data['YYMM'] = pd.to_datetime(data['YYMM'])  # 替代apply,更高效
processed_dfs = []  # 用列表存储每个处理后的子集

# 直接用groupby按YYMM分组处理
for date, group in data.groupby('YYMM'):
    df1 = group.sort_values(by='col1', ascending=False)
    df1['col2'] = df1['col1'] / df1['col1'].sum()
    df1['col3'] = reweight(df1['col2'], cap)
    processed_dfs.append(df1)

# 一次性合并所有子集,性能远高于循环append
data1 = pd.concat(processed_dfs, ignore_index=True)

方案2:用groupby.apply进一步简化代码

可以把分组处理逻辑封装成函数,直接用groupby.apply,代码更简洁高效:

def process_group(group, cap):
    group_sorted = group.sort_values(by='col1', ascending=False)
    group_sorted['col2'] = group_sorted['col1'] / group_sorted['col1'].sum()
    group_sorted['col3'] = reweight(group_sorted['col2'], cap)
    return group_sorted

data = pd.read_csv("df.csv")
data['YYMM'] = pd.to_datetime(data['YYMM'])
data1 = data.groupby('YYMM').apply(process_group, cap=cap).reset_index(drop=True)

关键优化点说明

  • 替换append为concat:列表存储+一次性concat避免了反复创建新DataFrame的开销,性能提升非常明显。
  • 给reweight加保护机制:最大迭代次数+浮点数误差容忍,彻底避免死循环,同时保留异常提示便于排查。
  • 优化日期转换:用pd.to_datetime(data['YYMM'])替代apply(pd.to_datetime),底层实现更高效。
  • 用groupby直接分组:避免了循环里反复筛选数据的开销,代码逻辑更清晰。

内容的提问来源于stack exchange,提问作者americ998

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.30 18:17:50