基于分组均值填充DataFrame列中的NaN值
按Hour+Location分组均值填充NaN的解决方案
你已经生成了分组均值的DataFrame,但填充失败大概率是因为索引未对齐或者没有针对每列匹配对应的分组均值。下面给两种直接可行的解决方法:
方法一:一步到位(推荐)
不用单独生成df_mean,直接用groupby.transform对目标列做填充,自动匹配分组均值:
# 先筛选需要填充的数值列(排除Hour、Location) numeric_cols = df.select_dtypes(include=['float64', 'int64']).columns.drop(['Hour', 'Location']) # 按Hour+Location分组,对每个数值列用组内均值填充NaN df[numeric_cols] = df.groupby(['Hour', 'Location'])[numeric_cols].transform( lambda x: x.fillna(x.mean(numeric_only=True)) )
这个方法会自动为每个NaN找到对应的Hour+Location组的均值,完全不用手动对齐索引,效率也更高。
方法二:利用已生成的df_mean填充
如果你一定要用已经生成的df_mean,需要先处理索引对齐问题,比如把df_mean的多级索引转成普通列,再合并到原DataFrame后填充:
# 把df_mean的多级索引(Hour、Location)转成普通列 df_mean_reset = df_mean.reset_index() # 合并原df和df_mean_reset,按Hour+Location匹配 df_merged = df.merge(df_mean_reset, on=['Hour', 'Location'], suffixes=('', '_mean')) # 遍历每个数值列,用_mean列填充原列的NaN for col in numeric_cols: df_merged[col] = df_merged[col].fillna(df_merged[f'{col}_mean']) # 清理临时生成的_mean列,得到最终结果 df = df_merged.drop([f'{col}_mean' for col in numeric_cols], axis=1)
为什么你之前的尝试失败?
df.groupby(['Hour','Location']).mean()生成的是**多级索引(Hour, Location)**的DataFrame,直接用df.fillna(df_mean)会因为索引不匹配导致填充无效——原df的普通索引和df_mean的多级索引无法对应,所以NaN不会被替换。
内容的提问来源于stack exchange,提问作者user26915184
相关产品推荐
相关产品推荐

