You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何消除Python中因循环添加列致DataFrame碎片化的性能警告?

解决DataFrame碎片化的PerformanceWarning警告

问题背景

我在Databricks中运行以下Python代码处理月度行业就业数据的DataFrame,逻辑为:追加空DataFrame、计算各行业就业的同比变化率、将每个序列前移1至12个月并将新生成的序列作为新列添加至原DataFrame:

# Forecast df parameters: index, column names
index = pd.date_range('2022-11-30', periods=13, freq='M')
columns = summary_table_empl.columns.to_list()                    

# Append history + empty df 
df_forecast = pd.DataFrame(index=index, columns=columns)
df_test_empl=pd.concat([summary_table_empl, df_forecast])

# New df, calculate yoy percent change for every commodity (col) 
df_CA_empl_test_yoy= ((df_test_empl - df_test_empl.shift(12))/df_test_empl.shift(12))*100

# Extend each variable as a series from 1 to 12 months
for col in df_CA_empl_test_yoy.columns:
   for i in range(1,13):
      df_CA_empl_test_yoy["%s_%s"%(col,i)] = df_CA_empl_test_yoy[col].shift(i)

df_CA_empl_test_yoy.head(15)

代码可正常运行,但触发如下警告:

PerformanceWarning: DataFrame is highly fragmented. This is usually
the result of calling frame.insert many times, which has poor
performance. Consider joining all columns at once using
pd.concat(axis=1) instead. To get a de-fragmented frame, use newframe = frame.copy() df_CA_empl_test_yoy["%s_%s"%(col,i)] = df_CA_empl_test_yoy[col].shift(i)

解决方案

优化思路

避免在循环中逐个向DataFrame添加列(每次添加都会触发内存碎片化),改为先一次性生成所有需要的移位列,再通过pd.concat完成合并,仅执行一次合并操作。

修改后的代码

# Forecast df parameters: index, column names
index = pd.date_range('2022-11-30', periods=13, freq='M')
columns = summary_table_empl.columns.to_list()                    

# Append history + empty df 
df_forecast = pd.DataFrame(index=index, columns=columns)
df_test_empl = pd.concat([summary_table_empl, df_forecast])

# New df, calculate yoy percent change for every commodity (col) 
df_CA_empl_test_yoy = ((df_test_empl - df_test_empl.shift(12))/df_test_empl.shift(12))*100

# 预先收集所有移位后的列,避免循环插入导致碎片化
shifted_cols = []
for col in df_CA_empl_test_yoy.columns:
    for i in range(1, 13):
        # 重命名Series以匹配原命名规则
        shifted_series = df_CA_empl_test_yoy[col].shift(i).rename(f"{col}_{i}")
        shifted_cols.append(shifted_series)

# 一次性合并所有新列到原DataFrame
df_CA_empl_test_yoy = pd.concat([df_CA_empl_test_yoy] + shifted_cols, axis=1)

df_CA_empl_test_yoy.head(15)

备选快速修复

如果不想重构代码,也可以在循环结束后执行以下语句,强制重新整理DataFrame内存结构以消除警告:

df_CA_empl_test_yoy = df_CA_empl_test_yoy.copy()

不过这种方法仅解决警告,性能提升不如第一种方案明显。

内容的提问来源于stack exchange,提问作者jack homareau

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 03:00:55