You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas按前两列同位置均值填充DataFrame缺失值

Pandas DataFrame 自定义规则缺失值填充实现

填充规则

  • 待处理对象为5列存在缺失值的DataFrame
  • 通用规则:所有列的NaN值,使用同一行内位于该列前方相邻两列的数值平均值替换
  • 特殊规则:第二列出现的第一个NaN值,取第一列的最后一个值填充

测试数据构造

使用如下代码生成初始测试DataFrame:

import pandas as pd

coh0 = [0.5, 0.3, 0.1, 0.2,0.2] 
coh1 = [0.4,0.3,0.6,0.5]
coh2 = [0.2,0.2,0.3]
coh3 = [0.8,0.8]
coh4 = [0.5]

df= pd.DataFrame({
    'coh0': pd.Series(coh0), 
    'coh1': pd.Series(coh1),
    'coh2': pd.Series(coh2), 
    'coh3': pd.Series(coh3),
    'coh4': pd.Series(coh4)
})

初始DataFrame内容:

coh0  coh1  coh2  coh3  coh4
0   0.5   0.4   0.2   0.8   0.5
1   0.3   0.3   0.2   0.8   NaN
2   0.1   0.6   0.3   NaN   NaN
3   0.2   0.5   NaN   NaN   NaN
4   0.2   NaN   NaN   NaN   NaN

实现代码

按照规则从左到右逐列填充即可,先处理第二列的特殊填充逻辑,再依次处理后续列:

# 处理第二列(coh1)的特殊填充规则:第一个NaN取第一列(coh0)最后一个值
coh1_first_nan_pos = df[df['coh1'].isna()].index[0]
df.loc[coh1_first_nan_pos, 'coh1'] = df['coh0'].iloc[-1]

# 从第三列开始逐列遍历,用同一行前两列均值填充NaN
col_list = df.columns.tolist()
for col_idx in range(2, len(col_list)):
    current_col = col_list[col_idx]
    nan_row_indexes = df[df[current_col].isna()].index
    for row_idx in nan_row_indexes:
        prev1_val = df.loc[row_idx, col_list[col_idx-1]]
        prev2_val = df.loc[row_idx, col_list[col_idx-2]]
        df.loc[row_idx, current_col] = (prev1_val + prev2_val) / 2

填充结果

运行代码后最终得到的DataFrame如下,完全符合填充规则要求:

coh0  coh1  coh2   coh3     coh4
0   0.5   0.4  0.20  0.800  0.50000
1   0.3   0.3  0.20  0.800  0.50000
2   0.1   0.6  0.30  0.450  0.37500
3   0.2   0.5  0.35  0.425  0.38750
4   0.2   0.2  0.20  0.200  0.20000

内容的提问来源于stack exchange,提问作者Debasis

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.30 14:57:28