R语言基于分组多行条件为数据框创建指定新列的实现问题
R分组生成条件列解决方案
方案1:dplyr实现(直观易读)
完整可运行代码:
# 构造原始数据框 df <- data.frame( Name = c("AI147", "AI147", "AI147", "AI147", "AI147", "AI20", "AI20", "AI87", "AI88", "AI88", "AI88", "AI65", "AI65"), Presence1 = c("both_type1", "soil", "soil", "water", "both_type2", "soil", "water", "both_type2", "soil", "soil", "soil", "water", "water") ) # 加载dplyr包 library(dplyr) df_result <- df %>% group_by(Name) %>% mutate(Final = case_when( # 匹配规则1:唯一值超过1种,或存在both_type前缀的取值 n_distinct(Presence1) > 1 | any(startsWith(Presence1, "both_type")) ~ "both", # 匹配规则2:组内全为soil all(Presence1 == "soil") ~ "soil", # 匹配规则3:组内全为water all(Presence1 == "water") ~ "water" )) %>% ungroup()
方案2:基础R实现(无需额外安装包)
完整可运行代码:
# 构造原始数据框同上 df <- data.frame( Name = c("AI147", "AI147", "AI147", "AI147", "AI147", "AI20", "AI20", "AI87", "AI88", "AI88", "AI88", "AI65", "AI65"), Presence1 = c("both_type1", "soil", "soil", "water", "both_type2", "soil", "water", "both_type2", "soil", "soil", "soil", "water", "water") ) # 用ave函数分组计算生成Final列 df$Final <- ave(df$Presence1, df$Name, FUN = function(group_val) { if (length(unique(group_val)) > 1 || any(startsWith(group_val, "both_type"))) { return("both") } else if (all(group_val == "soil")) { return("soil") } else if (all(group_val == "water")) { return("water") } })
两种方案运行后输出结果与你提供的预期结构完全一致。
内容的提问来源于stack exchange,提问作者ISalvi
相关产品推荐
相关产品推荐

