Pandas DataFrame碎片化警告解析及优化疑问
问题背景
我有一个包含111列、75000行的Pandas DataFrame,希望基于其他列的计算添加6个新列,初始代码如下:
df['col1_4'] = df['col1_1'] - df['col1_2'] - df['col1_3'] df['col2_4'] = df['col2_1'] - df['col2_2'] - df['col2_3'] df['col3_4'] = df['col3_1'] - df['col3_2'] - df['col3_3'] df['col4_4'] = df['col4_1'] - df['col4_2'] - df['col4_3'] df['col5_4'] = df['col5_1'] - df['col5_2'] - df['col5_3'] df['col6_4'] = df['col6_1'] - df['col6_2'] - df['col6_3']
执行后收到如下警告:
PerformanceWarning: DataFrame is highly fragmented. This is usually the result of calling
frame.insertmany times, which has poor performance. Consider joining all columns at once using pd.concat(axis=1) instead. To get a de-fragmented frame, usenewframe = frame.copy()
随后我将代码改写为一次性通过pd.concat添加列,执行后未收到警告,但仍有疑问:
- 是否需要警惕碎片化的DataFrame?其含义是什么?会对后续使用造成什么问题?
- 复制DataFrame为何能使其“去碎片化”?
改写后的代码:
col1 = df['col1_1'] - df['col1_2'] - df['col1_3'] col2 = df['col2_1'] - df['col2_2'] - df['col2_3'] col3 = df['col3_1'] - df['col3_2'] - df['col3_3'] col4 = df['col4_1'] - df['col4_2'] - df['col4_3'] col5 = df['col5_1'] - df['col5_2'] - df['col5_3'] col6 = df['col6_1'] - df['col6_2'] - df['col6_3'] df = pd.concat([df, col1, col2, col3, col4, col5, col6], axis=1)
问题解答
1. 碎片化DataFrame的含义、风险及是否需要警惕
Pandas的DataFrame底层由多个独立的数组(对应每一列)组成。当你多次单独添加列时,每次都会为新列分配一块独立的内存空间,这些新列的内存块和原有列的内存块大概率是不连续的——这就是DataFrame碎片化。
必须警惕这种情况,主要影响包括:
- 性能显著下降:后续对DataFrame的筛选、聚合、遍历等操作,需要在分散的内存块之间来回读取,内存访问效率大幅降低,数据量越大(比如你的7.5万行数据集),卡顿感越明显。
- 内存浪费:碎片化的内存布局会产生大量无法被高效利用的内存碎片,降低系统内存利用率。
- 潜在不稳定风险:极端情况下,频繁碎片化操作可能触发内存管理相关的异常,增加调试难度。
2. 复制DataFrame为何能去碎片化
当你调用newframe = frame.copy()时,Pandas会创建一个全新的DataFrame,并将所有列的数据连续地重新分配到一块完整的内存区域。原来分散在各个独立内存块的列数据,会被整理成连续的内存布局,相当于把零散的文件重新归档到一个连续的文件夹里,自然解决了碎片化问题。
为什么用pd.concat一次性添加列能解决警告
改写后的代码先计算出所有新列,再通过pd.concat一次性合并到原DataFrame中。这个过程中,Pandas会一次性为所有新列分配连续的内存空间,再和原DataFrame的内存块整合,避免了多次单独添加列导致的内存碎片化,因此不会触发PerformanceWarning。
内容的提问来源于stack exchange,提问作者Samba

