You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

能否用np.where实现Pandas DataFrame的NaN值前序有效值填充?

问题:使用np.where填充Pandas DataFrame的NaN值为最近前置有效值?

现有包含NaN值的Pandas DataFrame:

result
0       1
1      NaN
2      NaN
3       1
4      NaN
5       2
6       2
7      NaN
8       1

需求:将所有NaN值替换为其之前最近的有效值,预期结果如下:

result  Expected_Result
0       1            1
1      NaN           1
2      NaN           1
3       1            1
4      NaN           1
5       2            2
6       2            2
7      NaN           2
8       1            1

提问:是否可以使用np.where实现该填充操作?


回答

可以用np.where实现这个填充需求,但得配合一点辅助逻辑——毕竟np.where只是做条件取值,没法直接自动追踪最近的前置有效值,步骤如下:

  1. 先标记出所有非NaN的位置,然后生成一个索引数组,把NaN的位置替换成前一个有效数据的索引:
import pandas as pd
import numpy as np

# 构造原始DataFrame
df = pd.DataFrame({'result': [1, np.nan, np.nan, 1, np.nan, 2, 2, np.nan, 1]})

# 标记哪些位置是有效值
valid_mask = df['result'].notna()
# 创建数组:有效位置保留自身索引,NaN位置先设为nan
valid_indices = np.where(valid_mask, np.arange(len(df)), np.nan)
# 向前填充这个索引数组,让每个NaN位置拿到最近的有效索引
valid_indices = pd.Series(valid_indices).ffill().astype(int)
# 用np.where判断:是有效值就留原数,否则取对应有效索引的值
df['filled_result'] = np.where(valid_mask, df['result'], df['result'].iloc[valid_indices].values)
  1. 运行后得到的结果和预期完全一致:
result  filled_result
0     1.0            1.0
1     NaN            1.0
2     NaN            1.0
3     1.0            1.0
4     NaN            1.0
5     2.0            2.0
6     2.0            2.0
7     NaN            2.0
8     1.0            1.0

不过要提一句,Pandas本身自带的ffill()方法(即df['result'].ffill())能直接实现这个需求,代码更简洁。用np.where属于手动实现类似逻辑,可行但没必要,除非你有特殊场景需要用np.where的条件框架。

内容的提问来源于stack exchange,提问作者user28082680

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.16 14:17:22