R语言如何按所属国家条件为教育变量设置不同的因子标签
问题排查
你写的代码存在3个核心问题:
if()只能处理单逻辑值,无法对整列的国家向量做逐行判断,运行时只会用第一行的国家值走对应分支,所有行的编码规则都会和第一行保持一致- 国家名称匹配错误:你构造测试数据时
Country列的取值是带空格的Country 1、Country 2,判断条件里写的是无空格的Country1、Country2,完全匹配不上,所有行都会走到第三个else分支 - 变量引用错误:你生成的教育编码列叫
Education_1,分支里调用的df$Education是还没创建的新列,运行会直接报错
修正方案
方案1:基础R实现(无需额外安装包)
# 生成测试数据 df <- data.frame( Country = sample(c("Country 1", "Country 2", "Country 3"), 100, replace = TRUE), Education_1 = sample(1:4, 100, replace = TRUE) # 补充replace=TRUE生成100个样本,避免长度不匹配 ) # 逐行匹配国家设置因子 df$Education <- factor( mapply(function(country, code) { if (country %in% c("Country 1", "Country 2")) { switch(as.character(code), "1" = "No education", "2" = "Primary", "3" = "Secondary", "4" = NA_character_) } else { switch(as.character(code), "1" = "No education", "2" = "Islamic education", "3" = "Primary", "4" = "Secondary") } }, df$Country, df$Education_1), # 统一设置全数据集的因子水平,方便后续分析 levels = c("No education", "Islamic education", "Primary", "Secondary") )
方案2:tidyverse实现(代码更简洁易读)
library(dplyr) df <- df %>% mutate(Education = case_when( Country %in% c("Country 1", "Country 2") & Education_1 == 1 ~ "No education", Country %in% c("Country 1", "Country 2") & Education_1 == 2 ~ "Primary", Country %in% c("Country 1", "Country 2") & Education_1 == 3 ~ "Secondary", Country == "Country 3" & Education_1 == 1 ~ "No education", Country == "Country 3" & Education_1 == 2 ~ "Islamic education", Country == "Country 3" & Education_1 == 3 ~ "Primary", Country == "Country 3" & Education_1 == 4 ~ "Secondary", TRUE ~ NA_character_ ) %>% factor(levels = c("No education", "Islamic education", "Primary", "Secondary")))
内容的提问来源于stack exchange,提问作者jbell403
相关产品推荐
相关产品推荐

