合并排除哑变量:重新归类种族变量的R代码问题求助
问题
我有一个包含8个种族变量的数据集,需要按特定条件调整变量:
- 受访者可多选种族,其中
ethnicity_2代表白人 - 要创建新的
white变量,仅标记那些只选了ethnicity_2、未选择其他任何种族的受访者 - 尝试的代码出现三个问题:
dnresp变量(标记未选任何种族的样本)被错误赋值给几乎所有样本- 部分同时选了
ethnicity_2和其他种族的受访者被错误标记为white - 用
{{ethnicities.19}}的写法时,无论判断==1还是==0,都得到所有受访者未选任何种族的错误结果
原尝试代码:
ethnicities.19 <- c("ethnicity_1", "ethnicity_2", "ethnicity_3", "ethnicity_4", "ethnicity_5", "ethnicity_6", "ethnicity_7", "ethnicity_8") bar <- foo %>% select(ID, ethnicity_1:ethnicity_8) %>% mutate(across(.cols=ethnicity_1:ethnicity_8, .fns=function(x) { ifelse(is.na(x), 0, x)} )) %>% rowwise() %>% mutate(dnresp=ifelse(sum(eval(as.name(ethnicities.19)))==0, 1, 0), ## dnresp=ifelse(!any(eval(as.name(ethnicities.19))==1), 1, 0), white=ifelse(eval(as.name(ethnicities.19[2]))==1 & sum(eval(as.name(ethnicities.19[-c(2)])))==0, 1, 0))
错误尝试的代码片段:
dnresp=ifelse(!any({{ethnicities.19}}==1), 1, 0))
dnresp=ifelse(!any({{ethnicities.19}}==0), 1, 0))
数据样本:
structure(list(ID = c("ATL_01", "ATL_02", "ATL_03", "ATL_04", "ATL_05", "ATL_06", "ATL_07", "ATL_08", "ATL_09", "ATL_10", "ATL_11", "ATL_12", "ATL_13", "ATL_14", "ATL_15", "ATL_16", "ATL_17", "ATL_18", "ATL_19", "ATL_20"), ethnicity_1 = c(NA_real_, NA_real_, NA_real_, NA_real_, NA_real_, NA_real_, NA_real_, NA_real_, NA_real_, NA_real_, NA_real_, NA_real_, NA_real_, NA_real_, NA_real_, NA_real_, NA_real_, NA_real_, NA_real_, NA_real_), ethnicity_2 = c(1, 1, 1, NA, 1, NA, 1, NA, NA, 1, NA, NA, 1, NA, NA, NA, NA, 1, NA, NA), ethnicity_3 = c(NA, NA, NA, 1, NA, 1, NA, 1, 1, NA, 1, 1, NA, 1, 1, 1, 1, NA, 1, 1), ethnicity_4 = c(NA_real_, NA_real_, NA_real_, NA_real_, NA_real_, NA_real_, NA_real_, NA_real_, NA_real_, NA_real_, NA_real_, NA_real_, NA_real_, NA_real_, NA_real_, NA_real_, NA_real_, NA_real_, NA_real_, NA_real_), ethnicity_5 = c(NA_real_, NA_real_, NA_real_, NA_real_, NA_real_, NA_real_, NA_real_, NA_real_, NA_real_, NA_real_, NA_real_, NA_real_, NA_real_, NA_real_, NA_real_, NA_real_, NA_real_, NA_real_, NA_real_, NA_real_), ethnicity_6 = c(NA_real_, NA_real_, NA_real_, NA_real_, NA_real_, NA_real_, NA_real_, NA_real_, NA_real_, NA_real_, NA_real_, NA_real_, NA_real_, NA_real_, NA_real_, NA_real_, NA_real_, NA_real_, NA_real_, NA_real_), ethnicity_7 = c(NA_real_, NA_real_, NA_real_, NA_real_, NA_real_, NA_real_, NA_real_, NA_real_, NA_real_, NA_real_, NA_real_, NA_real_, NA_real_, NA_real_, NA_real_, NA_real_, NA_real_, NA_real_, NA_real_, NA_real_), ethnicity_8 = c(NA_real_, NA_real_, NA_real_, NA_real_, NA_real_, NA_real_, NA_real_, NA_real_, NA_real_, NA_real_, NA_real_, NA_real_, NA_real_, NA_real_, NA_real_, NA_real_, NA_real_, NA_real_, NA_real_, NA_real_)), row.names = c(NA, -20L), class = c("tbl_df", "tbl", "data.frame"))
问题原因
原代码的核心问题是**eval(as.name())无法处理字符串向量**:
as.name(ethnicities.19)只会把向量的第一个元素转为符号,后续元素被忽略,导致sum()和any()的计算完全错误{{ethnicities.19}}是tidy eval语法,用于传递变量名而非字符串向量,直接使用逻辑不通
解决方案
不用rowwise(),改用rowSums()更高效准确,直接针对列进行计算:
# 定义种族变量名向量 ethnicities.19 <- c("ethnicity_1", "ethnicity_2", "ethnicity_3", "ethnicity_4", "ethnicity_5", "ethnicity_6", "ethnicity_7", "ethnicity_8") bar <- foo %>% select(ID, all_of(ethnicities.19)) %>% # 将NA替换为0 mutate(across(all_of(ethnicities.19), ~ifelse(is.na(.), 0, .))) %>% # 计算总选择数 mutate(total_selected = rowSums(across(all_of(ethnicities.19)))) %>% # 生成dnresp:未选任何种族的样本标记为1 mutate(dnresp = ifelse(total_selected == 0, 1, 0)) %>% # 生成white:仅选了ethnicity_2的样本标记为1 mutate(white = ifelse(ethnicity_2 == 1 & total_selected == 1, 1, 0)) %>% # 可选:不需要total_selected可删除 select(-total_selected)
验证结果
用提供的样本数据运行后:
white变量会标记ATL_01、ATL_02、ATL_03、ATL_05、ATL_07、ATL_10、ATL_13、ATL_18为1(这些样本仅选了ethnicity_2)dnresp变量所有样本都是0(样本中没有未选任何种族的记录)
替代方案(保留rowwise写法)
如果一定要用rowwise(),可以用c_across()获取行内的列值:
bar <- foo %>% select(ID, all_of(ethnicities.19)) %>% mutate(across(all_of(ethnicities.19), ~ifelse(is.na(.), 0, .))) %>% rowwise() %>% mutate( dnresp = ifelse(sum(c_across(all_of(ethnicities.19))) == 0, 1, 0), white = ifelse(ethnicity_2 == 1 & sum(c_across(all_of(ethnicities.19[-2]))) == 0, 1, 0) ) %>% ungroup()
内容的提问来源于stack exchange,提问作者Stuart
相关产品推荐
相关产品推荐

