You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在np.where()逻辑中保留NaN值而非将其转为0?

问题描述

我有如下格式的数据集:

id    case2_q6
0   300   3.0
1   304   4.0
2   306   3.0
3   309   1.0
4   311   3.0
5   312   4.0
6   314   NaN
7   315   2.0
8   316   3.0
9   317   3.0

我用以下np.where()代码生成新变量fluid_2:

df['fluid_2'] = np.where((df['case2_q6'] == 1) | (df['case2_q6'] == 2), 1, 0)

生成后的结果里,索引6的NaN被转为了0:

id    case2_q6  fluid_2
0   300   3.0       0
1   304   4.0       0
2   306   3.0       0
3   309   1.0       1
4   311   3.0       0
5   312   4.0       0
6   314   NaN       0
7   315   2.0       1
8   316   3.0       0
9   317   3.0       0

我希望fluid_2列能保留原数据中的NaN值,期望输出如下:

id    case2_q6  fluid_2
0   300   3.0       0
1   304   4.0       0
2   306   3.0       0
3   309   1.0       1
4   311   3.0       0
5   312   4.0       0
6   314   NaN       NaN
7   315   2.0       1
8   316   3.0       0
9   317   3.0       0
解决方案

方法1:嵌套np.where()优先处理NaN

在外层先判断case2_q6是否为NaN,若是则返回np.nan,否则执行原逻辑:

import numpy as np
df['fluid_2'] = np.where(df['case2_q6'].isna(), np.nan, 
                        np.where((df['case2_q6'] == 1) | (df['case2_q6'] == 2), 1, 0))

方法2:用pandas的mask()补全NaN

先按原逻辑生成fluid_2,再将case2_q6为NaN的位置替换为NaN:

df['fluid_2'] = np.where((df['case2_q6'] ==1) | (df['case2_q6'] ==2),1,0)
df['fluid_2'] = df['fluid_2'].mask(df['case2_q6'].isna())

也可以一步完成:

df['fluid_2'] = ((df['case2_q6'] ==1) | (df['case2_q6'] ==2)).astype(int).mask(df['case2_q6'].isna())

方法3:用np.select()多条件匹配

通过多条件列表明确处理优先级,优先匹配NaN的情况:

conditions = [
    df['case2_q6'].isna(),
    (df['case2_q6'] ==1) | (df['case2_q6'] ==2),
    True
]
choices = [np.nan, 1, 0]
df['fluid_2'] = np.select(conditions, choices)

内容的提问来源于stack exchange,提问作者hulio_entredas

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.27 04:57:34