如何按同类型同运营周期均值填充s_months、incidents列的nan值
方法1:使用你已计算完成的shipgroup填充
首先确保shipgroup以types和o_periods作为多级索引,没有的话先执行索引设置:shipgroup = shipgroup.set_index(['types', 'o_periods'])
之后直接匹配分组索引填充对应列的空值即可:
# 填充s_months列缺失值 ship['s_months'] = ship['s_months'].fillna( ship.set_index(['types', 'o_periods']).index.map(shipgroup['s_months']) ) # 填充incidents列缺失值 ship['incidents'] = ship['incidents'].fillna( ship.set_index(['types', 'o_periods']).index.map(shipgroup['incidents']) )
方法2:无需提前计算分组均值,一步完成填充
如果不需要单独留存分组均值结果,可以直接用groupby+transform实现需求,代码更简洁:
# 按分组均值填充s_months空值 ship['s_months'] = ship['s_months'].fillna( ship.groupby(['types', 'o_periods'])['s_months'].transform('mean') ) # 按分组均值填充incidents空值 ship['incidents'] = ship['incidents'].fillna( ship.groupby(['types', 'o_periods'])['incidents'].transform('mean') )
注意事项
如果存在某组types+o_periods下对应列全为空的情况,上述方法填充后该组的空值仍会保留,你可以根据需求在fillna后追加二次填充规则,比如填充全表均值:ship['s_months'] = ship['s_months'].fillna(ship['s_months'].mean())
内容的提问来源于stack exchange,提问作者Fateh Muhammad
相关产品推荐
相关产品推荐

