Python Pandas如何仅对同年同月数据执行fillna的bfill/ffill填充
问题背景
现有一个带时间索引的面板DataFrame,内部存在大量NaN缺失值,示例构造代码如下:
import pandas as pd import numpy as np np.random.seed(2021) dates = pd.date_range('20130226', periods=720) df = pd.DataFrame(np.random.randint(0, 100, size=(720, 3)), index=dates, columns=list('ABC')) for col in df.columns: df.loc[df.sample(frac=0.4).index, col] = pd.np.nan df
原始数据输出示例:
A B C 2013-02-26 NaN NaN NaN 2013-02-27 NaN NaN 44.0 2013-02-28 62.0 NaN 29.0 2013-03-01 21.0 NaN 24.0 2013-03-02 12.0 70.0 70.0 ... ... ... 2015-02-11 38.0 42.0 NaN 2015-02-12 67.0 NaN NaN 2015-02-13 27.0 10.0 74.0 2015-02-14 18.0 NaN NaN 2015-02-15 NaN NaN NaN
需求说明
仅对属于同一年、同一月的数据执行df.fillna(method='bfill')或者df.fillna(method='ffill')缺失值填充,不允许跨年月填充。执行bfill后的预期效果如下:
A B C 2013-02-26 62.0 NaN 44.0 2013-02-27 62.0 NaN 44.0 2013-02-28 62.0 NaN 29.0 2013-03-01 21.0 70.0 24.0 2013-03-02 12.0 70.0 70.0 ... ... ... 2015-02-11 38.0 42.0 74.0 2015-02-12 67.0 10.0 74.0 2015-02-13 27.0 10.0 74.0 2015-02-14 18.0 NaN NaN 2015-02-15 NaN NaN NaN
实现方案
可以通过Pandas的groupby按年月分组后对每个分组单独执行填充操作,分组逻辑天然限制填充逻辑不会跨年月,完全符合需求:
# 后向填充bfill版本 df_filled = df.groupby([df.index.year, df.index.month], group_keys=False).apply(lambda x: x.bfill()) # 前向填充ffill版本 # df_filled = df.groupby([df.index.year, df.index.month], group_keys=False).apply(lambda x: x.ffill())
参数说明
- 用
df.index.year和df.index.month作为分组键,实现同一年同一月的数据被划分到同一个分组 group_keys=False用于避免分组后额外新增年月索引层级,保持输出DataFrame的索引结构和原始数据一致- 直接调用
bfill()/ffill()方法兼容Pandas新旧版本,避免旧版method参数弃用警告
内容的提问来源于stack exchange,提问作者ah bon
相关产品推荐
相关产品推荐

