DataFrame高度碎片化性能警告修复求助
问题背景
遇到如下性能警告:
PerformanceWarning: DataFrame is highly fragmented. This is usually the result of calling
frame.insertmany times, which has poor performance. Consider joining all columns at once using pd.concat(axis=1) instead. To get a de-fragmented frame, usenewframe = frame.copy()
触发警告的原代码:
df['xcount'] = df.apply(self.go_unigram, axis=1) df[self.listsunigram] = pd.DataFrame(df.xcount.tolist(), index=df.index) df['xcount'] = df.apply(self.go_bigram, axis=1) df[self.listsbigram] = pd.DataFrame(df.xcount.tolist(), index=df.index) df['xcount'] = df.apply(self.go_complex, axis=1) df[self.listcomplex] = pd.DataFrame(df.xcount.tolist(), index=df.index)
其中self.listsunigram/self.listsbigram/self.listcomplex是多列名列表,xcount是函数返回的多值列表,需求是将xcount的值分配到对应列中。
问题原因
原代码通过多次df[列列表] = ...的方式插入新列,每次赋值都会触发多次frame.insert操作,导致DataFrame内部存储碎片化,既触发警告,也拉低性能。
修复方案
核心思路是批量生成所有需要的新列,一次性合并到原DataFrame,避免多次插入操作:
# 分别生成三组新列对应的DataFrame unigram_cols = pd.DataFrame( df.apply(self.go_unigram, axis=1).tolist(), columns=self.listsunigram, index=df.index ) bigram_cols = pd.DataFrame( df.apply(self.go_bigram, axis=1).tolist(), columns=self.listsbigram, index=df.index ) complex_cols = pd.DataFrame( df.apply(self.go_complex, axis=1).tolist(), columns=self.listcomplex, index=df.index ) # 一次性合并所有新列到原DataFrame df = pd.concat([df, unigram_cols, bigram_cols, complex_cols], axis=1)
额外优化(可选)
如果self.go_unigram等函数支持向量化改写,建议替换掉apply(axis=1)——逐行处理的apply本身性能较低,改成向量化操作能进一步提升整体效率。
如果原DataFrame已经存在碎片化问题,可以执行df = df.copy()彻底整理存储结构,但上面的修复方案从源头避免了碎片化,通常无需额外执行此操作。
内容的提问来源于stack exchange,提问作者Paulo Alves

