使用pandas的DataFrame.loc插入数据时为何触发性能警告?
问题分析:DataFrame.loc插入数据的性能警告差异
现象复现
- 方式1(固定列标签、可变索引标签):循环添加行,无性能警告
import pandas as pd df=pd.DataFrame() for col in range(200): df.loc[str(col),'A']=str(col) - 方式2(固定索引标签、可变列标签):循环添加列,触发性能警告
df1=pd.DataFrame() for col in range(200): df1.loc['A',str(col)]=str(col)PerformanceWarning: DataFrame is highly fragmented. This is usually the result of calling
frame.insertmany times, which has poor performance. Consider joining all columns at once using pd.concat(axis=1) instead. To get a de-fragmented frame, usenewframe = frame.copy()
原因解析
Pandas的DataFrame采用列优先的内存存储模式:每一列是一段连续的内存数组。
- 方式1是给固定列
A逐行追加数据,本质是在现有列的内存块末尾扩展,内存分配高效,不会产生碎片。 - 方式2是逐列新增,每次都要为整个DataFrame重新分配内存(要给每一行都新增对应列的元素),循环200次后会导致内存空间碎片化,触发性能警告。
优化方案
方案1:提前构造数据字典,一次性生成DataFrame
直接批量构造列数据,避免循环插入列:
import pandas as pd # 构造包含所有列的字典,值为对应索引的列表 col_data = {str(col): [str(col)] for col in range(200)} df1 = pd.DataFrame(col_data, index=['A'])
方案2:用pd.concat批量合并列
先生成多个单列DataFrame,再一次性横向合并:
import pandas as pd dfs = [] for col in range(200): temp_df = pd.DataFrame({str(col): [str(col)]}, index=['A']) dfs.append(temp_df) df1 = pd.concat(dfs, axis=1)
临时修复碎片化DataFrame
如果已经出现警告,可通过复制操作整理内存:
df1 = df1.copy()
内容的提问来源于stack exchange,提问作者agile
相关产品推荐
相关产品推荐

