You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

DataFrame高度碎片化性能警告修复求助

修复DataFrame碎片化的PerformanceWarning问题

问题背景

遇到如下性能警告:

PerformanceWarning: DataFrame is highly fragmented. This is usually the result of calling frame.insert many times, which has poor performance. Consider joining all columns at once using pd.concat(axis=1) instead. To get a de-fragmented frame, use newframe = frame.copy()

触发警告的原代码:

df['xcount'] = df.apply(self.go_unigram, axis=1)
df[self.listsunigram] = pd.DataFrame(df.xcount.tolist(), index=df.index)

df['xcount'] = df.apply(self.go_bigram, axis=1)
df[self.listsbigram] = pd.DataFrame(df.xcount.tolist(), index=df.index)

df['xcount'] = df.apply(self.go_complex, axis=1)
df[self.listcomplex] = pd.DataFrame(df.xcount.tolist(), index=df.index)

其中self.listsunigram/self.listsbigram/self.listcomplex是多列名列表,xcount是函数返回的多值列表,需求是将xcount的值分配到对应列中。

问题原因

原代码通过多次df[列列表] = ...的方式插入新列,每次赋值都会触发多次frame.insert操作,导致DataFrame内部存储碎片化,既触发警告,也拉低性能。

修复方案

核心思路是批量生成所有需要的新列,一次性合并到原DataFrame,避免多次插入操作:

# 分别生成三组新列对应的DataFrame
unigram_cols = pd.DataFrame(
    df.apply(self.go_unigram, axis=1).tolist(),
    columns=self.listsunigram,
    index=df.index
)
bigram_cols = pd.DataFrame(
    df.apply(self.go_bigram, axis=1).tolist(),
    columns=self.listsbigram,
    index=df.index
)
complex_cols = pd.DataFrame(
    df.apply(self.go_complex, axis=1).tolist(),
    columns=self.listcomplex,
    index=df.index
)

# 一次性合并所有新列到原DataFrame
df = pd.concat([df, unigram_cols, bigram_cols, complex_cols], axis=1)

额外优化(可选)

如果self.go_unigram等函数支持向量化改写,建议替换掉apply(axis=1)——逐行处理的apply本身性能较低,改成向量化操作能进一步提升整体效率。

如果原DataFrame已经存在碎片化问题,可以执行df = df.copy()彻底整理存储结构,但上面的修复方案从源头避免了碎片化,通常无需额外执行此操作。

内容的提问来源于stack exchange,提问作者Paulo Alves

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.05 01:57:04