如何在np.where()逻辑中保留NaN值而非将其转为0?
问题描述
我有如下格式的数据集:
id case2_q6 0 300 3.0 1 304 4.0 2 306 3.0 3 309 1.0 4 311 3.0 5 312 4.0 6 314 NaN 7 315 2.0 8 316 3.0 9 317 3.0
我用以下np.where()代码生成新变量fluid_2:
df['fluid_2'] = np.where((df['case2_q6'] == 1) | (df['case2_q6'] == 2), 1, 0)
生成后的结果里,索引6的NaN被转为了0:
id case2_q6 fluid_2 0 300 3.0 0 1 304 4.0 0 2 306 3.0 0 3 309 1.0 1 4 311 3.0 0 5 312 4.0 0 6 314 NaN 0 7 315 2.0 1 8 316 3.0 0 9 317 3.0 0
我希望fluid_2列能保留原数据中的NaN值,期望输出如下:
id case2_q6 fluid_2 0 300 3.0 0 1 304 4.0 0 2 306 3.0 0 3 309 1.0 1 4 311 3.0 0 5 312 4.0 0 6 314 NaN NaN 7 315 2.0 1 8 316 3.0 0 9 317 3.0 0
解决方案
方法1:嵌套np.where()优先处理NaN
在外层先判断case2_q6是否为NaN,若是则返回np.nan,否则执行原逻辑:
import numpy as np df['fluid_2'] = np.where(df['case2_q6'].isna(), np.nan, np.where((df['case2_q6'] == 1) | (df['case2_q6'] == 2), 1, 0))
方法2:用pandas的mask()补全NaN
先按原逻辑生成fluid_2,再将case2_q6为NaN的位置替换为NaN:
df['fluid_2'] = np.where((df['case2_q6'] ==1) | (df['case2_q6'] ==2),1,0) df['fluid_2'] = df['fluid_2'].mask(df['case2_q6'].isna())
也可以一步完成:
df['fluid_2'] = ((df['case2_q6'] ==1) | (df['case2_q6'] ==2)).astype(int).mask(df['case2_q6'].isna())
方法3:用np.select()多条件匹配
通过多条件列表明确处理优先级,优先匹配NaN的情况:
conditions = [ df['case2_q6'].isna(), (df['case2_q6'] ==1) | (df['case2_q6'] ==2), True ] choices = [np.nan, 1, 0] df['fluid_2'] = np.select(conditions, choices)
内容的提问来源于stack exchange,提问作者hulio_entredas
相关产品推荐
相关产品推荐

