使用前后值的条件均值填充Pandas DataFrame中的NaN值
Pandas DataFrame 替换NaN为前后非NaN值的条件均值
原始数据
我们有如下DataFrame:
| col1 | col2 | col3 |
|---|---|---|
| 5 | 3 | 9 |
| NaN | 6 | NaN |
| NaN | 3 | 7 |
| 7 | 8 | 5 |
| NaN | 3 | NaN |
| 2 | 2 | 4 |
需求是将所有NaN替换为对应列中前后最近的非NaN值的均值,连续的NaN会使用同一组前后值的均值填充,示例结果如下:
| col1 | col2 | col3 |
|---|---|---|
| 5 | 3 | 9 |
| 6 | 6 | 8 |
| 6 | 3 | 7 |
| 7 | 8 | 5 |
| 4.5 | 3 | 4.5 |
| 2 | 2 | 4 |
实现代码
可以通过自定义函数对每列单独处理,找到每个NaN的前后非NaN值并计算均值:
import pandas as pd import numpy as np # 构造原始DataFrame df = pd.DataFrame({ 'col1': [5, np.nan, np.nan, 7, np.nan, 2], 'col2': [3, 6, 3, 8, 3, 2], 'col3': [9, np.nan, 7, 5, np.nan, 4] }) def fill_with_adjacent_mean(col): # 获取当前列所有非NaN的索引和对应值 non_nan_idx = col.dropna().index non_nan_vals = col.dropna().values # 遍历每个NaN的位置 for idx in col[col.isna()].index: # 找到左边最近的非NaN索引 left_idx = non_nan_idx[non_nan_idx < idx].max() if any(non_nan_idx < idx) else None # 找到右边最近的非NaN索引 right_idx = non_nan_idx[non_nan_idx > idx].min() if any(non_nan_idx > idx) else None # 根据前后值存在情况计算均值 if left_idx is not None and right_idx is not None: mean_val = (col[left_idx] + col[right_idx]) / 2 elif left_idx is not None: mean_val = col[left_idx] elif right_idx is not None: mean_val = col[right_idx] else: mean_val = np.nan col[idx] = mean_val return col # 应用函数到所有列 filled_df = df.apply(fill_with_adjacent_mean, axis=0) # 查看结果 print(filled_df)
代码说明
- 自定义函数
fill_with_adjacent_mean针对单列处理,先提取列内所有非NaN的位置和值 - 对每个NaN位置,分别找到前后最近的非NaN值,计算二者均值进行填充
- 如果NaN位于列的开头(无左侧非NaN值)或结尾(无右侧非NaN值),则直接使用单侧的非NaN值填充;若整列都是NaN则保持不变
- 通过
df.apply将函数批量应用到所有列,实现全DataFrame的NaN替换
内容的提问来源于stack exchange,提问作者heng xia
相关产品推荐
相关产品推荐

