Pandas中fillna与布尔掩码赋值的性能对比及差异原因探究
Pandas填充缺失值:fillna() vs 布尔掩码的性能对比
初始测试与核心问题
在第三列包含NaN的DataFrame中,我们对比两种缺失值填充方法的性能:
测试代码
import pandas as pd import numpy as np np.random.seed(100) nrows = 10000000 nnan = 25000 df = pd.DataFrame(np.random.uniform(0,250000,size=(nrows,3))) ind_row = np.random.randint(0,nrows,nnan) df.loc[ind_row, 2] = np.nan df1 = df.copy() %timeit df1[2] = df1[2].fillna(999) df1 = df.copy() %timeit df1[2].fillna(999) df1 = df.copy() %timeit df1.loc[df1[2].isna(),2] = 999
耗时结果
第一次测试:
35.1 ms ± 369 μs per loop (mean ± std. dev. of 7 runs, 10 loops each) 36.4 ms ± 331 μs per loop (mean ± std. dev. of 7 runs, 10 loops each) 6.23 ms ± 65.5 μs per loop (mean ± std. dev. of 7 runs, 100 loops each)
第二次测试:
35.9 ms ± 1.11 ms per loop (mean ± std. dev. of 7 runs, 10 loops each) 37.3 ms ± 438 μs per loop (mean ± std. dev. of 7 runs, 10 loops each) 6.41 ms ± 38.2 μs per loop (mean ± std. dev. of 7 runs, 100 loops each)
测试发现耗时与NaN占比关联不大,核心问题:为何手动布尔掩码的性能优于fillna()?
更新1:不同DataFrame规模的耗时对比
测试结果图表显示:当行数≥10^5时,布尔掩码的耗时始终低于fillna(inplace=True)与普通fillna()。
测试代码
import pandas as pd import numpy as np import matplotlib.pyplot as plt import timeit np.random.seed(100) nrepeats = 5 nruns = 10 def time_fillna(nrows, output): nnan = int(nrows/4) df = pd.DataFrame(np.random.uniform(0,250000,size=(nrows,3))) ind_row = np.random.randint(0,nrows,nnan) df.loc[ind_row, 2] = np.nan df1 = df.copy() output['fillna_assign'].append(np.min(timeit.repeat('df1[2] = df1[2].fillna(999)', globals=locals(), repeat=nrepeats, number=nruns))/nruns) df1 = df.copy() output['fillna_only'].append(np.min(timeit.repeat('df1[2].fillna(999)', globals=locals(), repeat=nrepeats, number=nruns))/nruns) df1 = df.copy() output['fillna_inplace'].append(np.min(timeit.repeat('df1[2].fillna(999,inplace=True)', globals=locals(), repeat=nrepeats, number=nruns))/nruns) df1 = df.copy() output['bool_mask'].append(np.min(timeit.repeat('df1.loc[df1[2].isna(),2] = 999', globals=locals(), repeat=nrepeats, number=nruns))/nruns) output = {method:[] for method in ['fillna_assign', 'fillna_only', 'fillna_inplace', 'bool_mask']} nrows_all = np.logspace(3,8,11).astype(int) for nrows in nrows_all: print(f'Timing for nrows = {nrows}') time_fillna(nrows, output) plt.ion() plt.plot(nrows_all, output['fillna_assign'], label='fillna_assign') plt.plot(nrows_all, output['fillna_only'], label='fillna_only') plt.plot(nrows_all, output['fillna_inplace'], label='fillna_inplace') plt.plot(nrows_all, output['bool_mask'], label='bool_mask') plt.loglog() plt.xlabel('Number of rows') plt.ylabel('Runtime [s]') plt.legend() plt.show() plt.savefig('timing_fillna.png')
更新2:修正测试逻辑后的结果
修正测试逻辑(确保每次测试都基于全新的含NaN数据)后,结论发生变化:fillna(inplace=True)与布尔掩码性能相当,无赋值操作的fillna()速度最快。
修正后的测试代码
def time_fillna_revised(nrows, output): nnan = int(nrows/4) df = pd.DataFrame(np.random.uniform(0,250000,size=(nrows,3))) ind_row = np.random.randint(0,nrows,nnan) df.loc[ind_row, 2] = np.nan df_insert = df.copy() t_fillna_assign = [] t_fillna_inplace = [] t_fillna_bool = [] for i in range(nrepeats): time_temp_assign = 0 time_temp_inplace = 0 time_temp_bool = 0 for j in range(nruns): df1 = df.copy() time_temp_assign += timeit.repeat('df_insert[2] = df1[2].fillna(999)', globals=locals(), repeat=1, number=1)[0] df1 = df.copy() time_temp_inplace += timeit.repeat('df1[2].fillna(999,inplace=True)', globals=locals(), repeat=1, number=1)[0] df1 = df.copy() time_temp_bool += timeit.repeat('df1.loc[df1[2].isna(),2] = 999', globals=locals(), repeat=1, number=1)[0] t_fillna_assign.append(time_temp_assign) t_fillna_inplace.append(time_temp_inplace) t_fillna_bool.append(time_temp_bool) output['fillna_assign'].append(np.min(t_fillna_assign)/nruns) output['fillna_inplace'].append(np.min(t_fillna_inplace)/nruns) output['bool_mask'].append(np.min(t_fillna_bool)/nruns) df1 = df.copy() output['fillna_only'].append(np.min(timeit.repeat('df1[2].fillna(999)', globals=locals(), repeat=nrepeats, number=nruns))/nruns)
更新3:不同NaN占比的性能表现
- 当NaN占比为1/4时,性能结果更复杂,各方法的耗时差异随DataFrame规模变化出现交叉
- 当NaN占比为1/400时,布尔掩码的耗时显著低于fillna的各类调用方式
运行@mozway代码的结果
测试结果显示,fillna(inplace=True)与布尔掩码的耗时几乎重合,无赋值的fillna()依然是最快的选项。
性能差异的原因分析
- 通用逻辑开销:
fillna()是通用函数,需支持前向/后向填充、均值填充等多种策略,内部包含大量分支判断、参数校验和类型兼容逻辑,即使是简单常量填充,也会执行这些通用流程,带来额外开销。 - 数据复制差异:默认
fillna()会返回新的Series/DataFrame,涉及完整数据复制;即使使用inplace=True,仍存在初始化和校验步骤,而布尔掩码是直接定位NaN位置修改原数据,逻辑更直接。 - 操作粒度影响:当NaN占比较低时,布尔掩码仅对少量元素赋值,而
fillna()会遍历整个Series处理所有元素,效率更低;当NaN占比较高时,fillna()的批量处理优化会缩小与布尔掩码的差距,甚至反超。
内容的提问来源于stack exchange,提问作者silence_of_the_lambdas
相关产品推荐
相关产品推荐

