minimize()迭代间隔时长持续增加问题排查求助
我在排查优化过程耗时过长的问题时,给代码加了时间戳,发现上一次迭代return前的时间戳和当前迭代“start iteration”开始的时间间隔在持续增长,但函数内部各操作的耗时基本稳定。已经分析过内存,数据帧构建完成后内存没有波动,求排查思路。
相关代码如下:
import time import numpy as np from scipy.optimize import minimize def objective(df_initial_guess, shape, now): print("start iteration -", time.time()-now) now = time.time() df_initial_guess = df_initial_guess.reshape(shape) global df_training_data # 使用全局变量 print("initial guess and training data organized -", time.time()-now) now = time.time() # 省略中间代码...... print("formula spectrum calculation done -", time.time()-now) now = time.time() data=np.array(data)*100 df_training_data[57]=data print("anomaly calculation done -", time.time()-now) now = time.time() print() return df_training_data.loc[df_training_data.iloc[:,7] >= 1, 57].mean() # 假设df_initial_guess、shape、now已定义 result=minimize(objective, df_initial_guess, args=(shape, now,))
终端输出:
start iteration - 0.001001119613647461 initial guess and training data organized - 0.0 Pure color's constants were calculated - 0.2595968246459961 formulas' constants were calculated - 10.22897744178772 formula spectrum calculation done - 160.1121380329132 anomaly calculation done - 0.701399564743042 start iteration - 171.34229683876038 initial guess and training data organized - 0.015207767486572266 Pure color's constants were calculated - 0.18435430526733398 formulas' constants were calculated - 10.336299896240234 formula spectrum calculation done - 145.85290336608887 anomaly calculation done - 1.0375640392303467 start iteration - 328.769579410553 initial guess and training data organized - 0.0027430057525634766 Pure color's constants were calculated - 0.19941067695617676 formulas' constants were calculated - 14.058949947357178 formula spectrum calculation done - 133.65063190460205 anomaly calculation done - 0.6168942451477051 start iteration - 477.33042335510254 initial guess and training data organized - 0.00693202018737793 Pure color's constants were calculated - 0.1299881935119629 formulas' constants were calculated - 8.694848537445068 formula spectrum calculation done - 134.9823079109192 anomaly calculation done - 0.5860121250152588 start iteration - 621.7472784519196 initial guess and training data organized - 0.009764671325683594 Pure color's constants were calculated - 0.12981104850769043 formulas' constants were calculated - 8.949415445327759 formula spectrum calculation done - 146.67141890525818 anomaly calculation done - 0.46207499504089355 start iteration - 777.9782621860504 initial guess and training data organized - 0.0010035037994384766 Pure color's constants were calculated - 0.1138761043548584 formulas' constants were calculated - 7.141656398773193 formula spectrum calculation done - 114.04382729530334 anomaly calculation done - 0.48108625411987305 start iteration - 899.7692215442657 initial guess and training data organized - 0.0010166168212890625 Pure color's constants were calculated - 0.11710047721862793 formulas' constants were calculated - 7.446816444396973
排查优化器内部计算开销:你使用的
minimize(推测是scipy的实现)在迭代间隙可能执行梯度估计、线搜索、收敛性检查等操作。比如BFGS类算法需要维护Hessian近似矩阵,随着迭代次数增加,矩阵相关计算的耗时会累积。可以换用Nelder-Mead这类无梯度、内部维护成本低的优化器测试,看间隔耗时是否依然增长,以此确认是否是优化器本身的问题。检查全局DataFrame的隐性开销:虽然内存无波动,但每次迭代修改
df_training_data[57]=data时,Pandas可能会做索引维护、类型校验或缓存更新,这些操作的耗时不会直接体现在内存变化上,但会随着迭代次数累积。可以尝试用Numpy数组单独维护这一列,或者预先分配好空间直接赋值,对比测试耗时变化。监控系统层面的资源消耗:用
psutil库在迭代间隙监控CPU使用率、磁盘IO、内存交换情况,排查是否有后台进程抢占资源,或者Python垃圾回收(GC)在该时间段触发。可以手动关闭GC(gc.disable())测试,或者打印GC统计信息,确认是否是GC导致的耗时增长。检查优化器参数配置:查看
minimize的options参数,比如gtol(梯度收敛阈值)、maxiter等,是否设置了导致优化器在迭代间隙做更多收敛性检查的参数。如果使用的是需要梯度的优化器且未提供自定义梯度函数,优化器会用数值微分,后续迭代的数值计算量可能因参数变化或误差累积而增加。可以实现自定义梯度函数,对比耗时变化。验证返回值计算的隐性耗时:函数返回的
df_training_data.loc[...].mean(),虽然你在函数内做了计时,但优化器接收返回值后可能有额外处理,或者这行代码的实际耗时未被完全统计。可以提前计算并缓存结果,或者拆分计算步骤单独计时,确认是否这部分耗时被计入了迭代间隙。
内容的提问来源于stack exchange,提问作者sivan

