You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用pandas的DataFrame.loc插入数据时为何触发性能警告?

问题分析:DataFrame.loc插入数据的性能警告差异

现象复现

  • 方式1(固定列标签、可变索引标签):循环添加行,无性能警告
    import pandas as pd
    
    df=pd.DataFrame()
    
    for col in range(200):
        df.loc[str(col),'A']=str(col)
    
  • 方式2(固定索引标签、可变列标签):循环添加列,触发性能警告
    df1=pd.DataFrame()
    for col in range(200):
        df1.loc['A',str(col)]=str(col)
    

    PerformanceWarning: DataFrame is highly fragmented. This is usually the result of calling frame.insert many times, which has poor performance. Consider joining all columns at once using pd.concat(axis=1) instead. To get a de-fragmented frame, use newframe = frame.copy()

原因解析

Pandas的DataFrame采用列优先的内存存储模式:每一列是一段连续的内存数组。

  • 方式1是给固定列A逐行追加数据,本质是在现有列的内存块末尾扩展,内存分配高效,不会产生碎片。
  • 方式2是逐列新增,每次都要为整个DataFrame重新分配内存(要给每一行都新增对应列的元素),循环200次后会导致内存空间碎片化,触发性能警告。

优化方案

方案1:提前构造数据字典,一次性生成DataFrame

直接批量构造列数据,避免循环插入列:

import pandas as pd

# 构造包含所有列的字典,值为对应索引的列表
col_data = {str(col): [str(col)] for col in range(200)}
df1 = pd.DataFrame(col_data, index=['A'])

方案2:用pd.concat批量合并列

先生成多个单列DataFrame,再一次性横向合并:

import pandas as pd

dfs = []
for col in range(200):
    temp_df = pd.DataFrame({str(col): [str(col)]}, index=['A'])
    dfs.append(temp_df)
df1 = pd.concat(dfs, axis=1)

临时修复碎片化DataFrame

如果已经出现警告,可通过复制操作整理内存:

df1 = df1.copy()

内容的提问来源于stack exchange,提问作者agile

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.19 19:31:02