如何用同行相邻列均值替换DataFrame中的NaN值?
解决DataFrame中NaN值替换为相邻列均值的问题
示例数据
你的示例DataFrame如下:
import pandas as pd import numpy as np d = {'col1': [4, 0, 2], 'col2': [1, np.nan, 4], 'col3': [12, -2, 4]} df = pd.DataFrame(data=d)
方案1:通用处理(适用于任意位置的NaN)
如果你的DataFrame中NaN可能出现在任意列(首列、中间列、末列),可以定义一个行处理函数,遍历每行的NaN并计算相邻有效值的均值填充:
def fill_nan_with_neighbor_mean(row): nan_cols = row[row.isna()].index for col in nan_cols: col_idx = row.index.get_loc(col) neighbors = [] # 检查左侧相邻列 if col_idx > 0: left_val = row[row.index[col_idx - 1]] if not pd.isna(left_val): neighbors.append(left_val) # 检查右侧相邻列 if col_idx < len(row.index) - 1: right_val = row[row.index[col_idx + 1]] if not pd.isna(right_val): neighbors.append(right_val) # 用相邻值的均值填充NaN if neighbors: row[col] = np.mean(neighbors) return row # 应用函数到每行 df_filled = df.apply(fill_nan_with_neighbor_mean, axis=1) print(df_filled)
运行后输出结果:
col1 col2 col3 0 4 1.0 12 1 0 -1.0 -2 2 2 4.0 4
方案2:针对已知位置的NaN简化处理
如果已经明确NaN只出现在中间列(比如示例中的col2),且左右列都有有效值,可以直接计算左右列的均值填充:
df['col2'] = df['col2'].fillna((df['col1'] + df['col3']) / 2)
这种方式代码更简洁,适合场景明确的情况。
内容的提问来源于stack exchange,提问作者Davi
相关产品推荐
相关产品推荐

