如何用Column2的值覆盖Column1的值(排除Column2的NA值)
需求说明
我有两个种族相关列Race1和Race2,Race2包含对Race1的部分修正值,但大部分为NA值。
输入示例
Race1 Race2 White American Indian White American Indian White Black Black NA Black NA White NA White NA
期望输出
Race American Indian American Indian Black Black Black White White
需要实现:用Race2的值覆盖Race1的值,同时忽略Race2中的NA值。
实现方案
1. R语言(基于dplyr包)
使用coalesce()函数,它会按参数顺序取第一个非NA值,刚好匹配需求:
library(dplyr) # 假设数据框名为df df <- df %>% mutate(Race = coalesce(Race2, Race1)) %>% select(Race) # 仅保留新生成的Race列
2. Python(基于Pandas库)
有两种常用方法:
方法一:combine_first()
import pandas as pd # 假设数据框名为df df['Race'] = df['Race2'].combine_first(df['Race1']) df = df[['Race']] # 筛选保留Race列
方法二:fillna()
import pandas as pd df['Race'] = df['Race2'].fillna(df['Race1']) df = df[['Race']]
内容的提问来源于stack exchange,提问作者user17073706
相关产品推荐
相关产品推荐

