R语言dataframe新增行后缺失计算变量的低侵入式补全方法咨询
解决方案
方案1:基础R实现(无额外依赖,最低侵入)
仅针对age_category为NA的行做条件计算,不会覆盖已有的正确记录,全程在原数据集上操作,无需拆分合并:
# 仅对缺失值行做赋值计算,不改动已有有效值 test$age_category[is.na(test$age_category)] <- ifelse( test$age[is.na(test$age_category)] >= 65, "65+", ifelse(test$age[is.na(test$age_category)] >= 30 & test$age[is.na(test$age_category)] < 65, "30 to 64", ifelse(test$age[is.na(test$age_category)] >= 18 & test$age[is.na(test$age_category)] < 30, "18 to 29", "under 18") ) )
运行后输出结果:
name age age_category 1 mina 66 65+ 2 alex 24 18 to 29 3 katie 19 18 to 29 4 eric 36 30 to 64
方案2:dplyr实现(代码可读性更高,适合tidyverse生态用户)
用coalesce优先保留已有非空值,仅对缺失值走case_when的条件计算,嵌套层级低,不容易写错判断逻辑:
library(dplyr) test <- test %>% mutate(age_category = coalesce( age_category, # 优先保留原字段的非空值 case_when( age >= 65 ~ "65+", age >= 30 & age < 65 ~ "30 to 64", age >= 18 & age < 30 ~ "18 to 29", TRUE ~ "under 18" ) ))
内容的提问来源于stack exchange,提问作者ajt10
相关产品推荐
相关产品推荐

